NVIDIA Nemotron 3 Ultra API Drops 39% on Output Tokens
What Changed
NVIDIA has quietly repriced the Nemotron 3 Ultra API. Input tokens dropped from $0.60 to $0.50 per 1M tokens, while output tokens fell harder — from $3.60 to $2.20 per 1M tokens, a -39% cut overall.
Why Output Pricing Is What Matters Here
For most real-world workloads — summarization, code generation, agentic loops — output tokens dominate your bill. A drop from $3.60 to $2.20 per 1M output tokens is not cosmetic. If you were running a pipeline that previously cost $100/day in output alone, the same workload now runs closer to $61.
The input cut from $0.60 to $0.50 is smaller in absolute terms and will matter mainly for retrieval-augmented or long-context workflows where large chunks of text are stuffed into every prompt.
Should You Switch or Re-Evaluate?
If you benchmarked Nemotron 3 Ultra previously and ruled it out on cost, the output repricing is worth a second look — especially for generation-heavy tasks. If you're already on a cheaper alternative, the gap may have narrowed but check quality tradeoffs before migrating.
See current pricing alongside alternatives on the Nemotron 3 Ultra — live specs & price history page.
Bottom Line
A 39% output price cut is a real reduction, not a rounding-error discount. Output-heavy use cases get the most benefit. Input-heavy use cases see a modest improvement. Worth re-running your cost estimates if this model was on your shortlist.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.