Back to Blog

NVIDIA Nemotron 3 Ultra API Drops 39% on Output Tokens

EzraJuly 23, 20261 min read
NVIDIA Nemotron 3 Ultra API Drops 39% on Output Tokens

What Changed

NVIDIA has quietly repriced the Nemotron 3 Ultra API. Input tokens dropped from $0.60 to $0.50 per 1M tokens, while output tokens fell harder — from $3.60 to $2.20 per 1M tokens, a -39% cut overall.

Why Output Pricing Is What Matters Here

For most real-world workloads — summarization, code generation, agentic loops — output tokens dominate your bill. A drop from $3.60 to $2.20 per 1M output tokens is not cosmetic. If you were running a pipeline that previously cost $100/day in output alone, the same workload now runs closer to $61.

The input cut from $0.60 to $0.50 is smaller in absolute terms and will matter mainly for retrieval-augmented or long-context workflows where large chunks of text are stuffed into every prompt.

Should You Switch or Re-Evaluate?

If you benchmarked Nemotron 3 Ultra previously and ruled it out on cost, the output repricing is worth a second look — especially for generation-heavy tasks. If you're already on a cheaper alternative, the gap may have narrowed but check quality tradeoffs before migrating.

See current pricing alongside alternatives on the Nemotron 3 Ultra — live specs & price history page.

Bottom Line

A 39% output price cut is a real reduction, not a rounding-error discount. Output-heavy use cases get the most benefit. Input-heavy use cases see a modest improvement. Worth re-running your cost estimates if this model was on your shortlist.

Ezra, Scout AI Team

E

Ezra

Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.

NVIDIA Nemotron 3 Ultra API Drops 39% on Output Tokens | AIToolScout