Back to Blog

NVIDIA Nemotron 3 Ultra API Prices Drop Up to 39%

EzraAugust 28, 20261 min read
NVIDIA Nemotron 3 Ultra API Prices Drop Up to 39%

What Changed

NVIDIA has quietly repriced Nemotron 3 Ultra. Input tokens dropped from $0.60 to $0.50 per 1M tokens, while output tokens fell from $3.60 to $2.20 per 1M tokens — the headline -39% cut applies to that output rate.

Why Output Pricing Is the One to Watch

For most real-world workloads — summarisation, code generation, multi-turn chat — output tokens dominate the bill. A drop from $3.60 to $2.20 per 1M output tokens isn't cosmetic. If your app generates, say, 50M output tokens a month, that's a monthly saving of $70. Scale that to 500M tokens and you're looking at $700/month back in your budget without changing a single line of code.

Input pricing is now at $0.50 per 1M tokens, which is competitive for a model in this class, though it was never the pain point.

Should You Switch or Re-Evaluate?

If you already use Nemotron 3 Ultra, rerun your cost projections — this cut likely improves your unit economics materially. If you've been sitting on competing models primarily because of price, this is a reasonable moment to benchmark it against your use case.

If you're evaluating alternatives, our comparison tool lets you stack Nemotron 3 Ultra against similar models side by side. Check the Nemotron 3 Ultra — live specs & price history page for up-to-date pricing and how it stacks up.

Bottom Line

A -39% output price cut is substantial, not incremental. For output-heavy production workloads, this changes the calculus enough to warrant a fresh look — even if you ruled out Nemotron 3 Ultra on cost grounds before.

Ezra, Scout AI Team

E

Ezra

Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.

NVIDIA Nemotron 3 Ultra API Prices Drop Up to 39% | AIToolScout