NVIDIA Nemotron 3.5 Lightning Listed: $0.10 Input, 262k Context
What Just Dropped
NVIDIA added Nemotron 3.5 Lightning to its model lineup on 2026-08-11. It's a text-to-text, open-source model priced at $0.10 per 1M input tokens and $0.25 per 1M output tokens, with a 262k-token context window.
Why the Pricing Catches Attention
At $0.10 input / $0.25 output, Nemotron 3.5 Lightning sits in the budget-to-mid tier — competitive territory that's increasingly crowded, but the 262k context is a genuine differentiator at this price point. Long-document summarization, large codebase analysis, or multi-turn agents that need to hold a lot in memory are the obvious use cases.
For a concrete sanity check: processing 10M input tokens costs $1.00 flat. Output-heavy workloads (think generative pipelines with verbose responses) will nudge costs up faster given the 2.5× markup between input and output rates — worth modeling before committing.
Open-Source Angle
Being open-source means you can self-host to eliminate per-token costs entirely if your infrastructure supports it. That changes the calculus significantly for high-volume users who can absorb fixed compute costs instead.
Should You Switch?
If you're currently paying more than $0.10/1M input on a comparable text model and don't need multimodal capabilities, Nemotron 3.5 Lightning is worth a benchmark run. If your workloads are short-context and output-light, the 262k window is wasted headroom — cheaper alternatives may still win on unit economics.
See full specs, context limits, and track price changes at Nemotron 3.5 Lightning — live specs & price history.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.