NVIDIA Nemotron 3 Ultra API Prices Drop Up to 39%
What Changed
NVIDIA has quietly repriced Nemotron 3 Ultra. Input tokens dropped from $0.60 to $0.50 per 1M tokens, while output tokens fell from $3.60 to $2.20 per 1M tokens — the headline -39% cut applies to that output rate.
Why Output Pricing Is the One to Watch
For most real-world workloads — summarisation, code generation, multi-turn chat — output tokens dominate the bill. A drop from $3.60 to $2.20 per 1M output tokens isn't cosmetic. If your app generates, say, 50M output tokens a month, that's a monthly saving of $70. Scale that to 500M tokens and you're looking at $700/month back in your budget without changing a single line of code.
Input pricing is now at $0.50 per 1M tokens, which is competitive for a model in this class, though it was never the pain point.
Should You Switch or Re-Evaluate?
If you already use Nemotron 3 Ultra, rerun your cost projections — this cut likely improves your unit economics materially. If you've been sitting on competing models primarily because of price, this is a reasonable moment to benchmark it against your use case.
If you're evaluating alternatives, our comparison tool lets you stack Nemotron 3 Ultra against similar models side by side. Check the Nemotron 3 Ultra — live specs & price history page for up-to-date pricing and how it stacks up.
Bottom Line
A -39% output price cut is substantial, not incremental. For output-heavy production workloads, this changes the calculus enough to warrant a fresh look — even if you ruled out Nemotron 3 Ultra on cost grounds before.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.