NVIDIA Nemotron 3 Ultra API Prices Cut: Output Tokens Drop 23%
What Changed
NVIDIA has reduced Nemotron 3 Ultra's API pricing across both token directions. Input costs nudged down from $0.625 to $0.60 per 1M tokens, while output costs dropped more sharply — from $3.13 to $2.40 per 1M tokens. The blended reduction works out to roughly -23%, though the real saving lands on the output side where most production costs accumulate.
Does It Matter for Your Workload?
For read-heavy or classification tasks where input tokens dominate, the $0.025 input saving per 1M tokens is marginal. But if you're running long-form generation, agentic loops, or RAG pipelines with verbose outputs, the drop from $3.13 to $2.40 on output tokens is worth recalculating. At moderate scale — say, 50M output tokens a month — that's roughly $365 back in your budget monthly without changing a line of code.
Nemotron 3 Ultra was already positioned as a capable reasoning model at a mid-tier price point. This cut makes it incrementally more competitive against similarly sized models, though developers should benchmark quality-per-dollar against alternatives before committing budgets.
Next Steps
Check the updated numbers, historical pricing, and how it stacks up against comparable models on the Nemotron 3 Ultra — live specs & price history page before your next infrastructure review.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.