NVIDIA Nemotron 3 Super API Drops 11% on Output Tokens
What Changed
NVIDIA has repriced the Nemotron 3 Super API. Output tokens drop from $0.45 to $0.40 per 1M tokens — an 11% reduction. Input tokens move in the opposite direction, edging up from $0.08 to $0.085 per 1M tokens.
The headline -11% figure applies to the output side only, so the real-world impact depends heavily on your input/output ratio.
Who Feels It Most
Workloads that are output-heavy — long-form generation, summarization, code completion, agentic pipelines — will see a meaningful cost improvement. At scale, shaving $0.05 per 1M output tokens adds up quickly: a workflow burning 100M output tokens monthly saves $5,000.
Conversely, if your use case is input-heavy (large context retrieval, document analysis with short responses), the small input price increase partially offsets the gain. The crossover point sits somewhere around a 10:1 input-to-output token ratio — above that, the net change may be negligible or slightly negative.
Worth Comparing
Before locking in, it's worth stacking this against alternatives on our site. Nemotron 3 Super is a 120B-parameter mixture-of-experts model activating 12B parameters per forward pass — that architecture is what keeps inference costs competitive in the first place. Whether the new pricing makes it the best value at this capability tier depends on your specific benchmark priorities.
See the full breakdown, live specs, and historical pricing at Nemotron 3 Super — live specs & price history.
Bottom Line
Output-heavy teams: this is a straightforward win. Input-heavy teams: run your token ratio numbers before calling it a cut. Either way, the repricing is worth a quick audit of your current provider costs.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.