NVIDIA Nemotron 3 Ultra API Prices Drop: Input Now $0.50, Output $2.20 per 1M Tokens
What Changed
NVIDIA has quietly trimmed API pricing on Nemotron 3 Ultra. Input tokens drop from $0.60 to $0.50 per 1M tokens, and output tokens fall from $2.40 to $2.20 per 1M tokens — a -17% reduction across the board.
Does It Move the Needle for Your Workload?
A 17% cut is meaningful but not dramatic. Where it matters most:
- Output-heavy workloads (long-form generation, summarisation pipelines) benefit most, since output tokens were already 4× the input rate and represent the bigger spend for most production apps.
- High-volume inference at scale — if you're pushing millions of requests, shaving $0.20 per 1M output tokens compounds quickly.
- Low-volume or prototype usage — the savings are real but unlikely to change a build/don't-build decision on their own.
The input-to-output price gap remains wide ($0.50 vs $2.20), so workloads that generate verbose responses still face a steep output bill relative to peers. Before committing, it's worth stacking Nemotron 3 Ultra against comparable models on our site — particularly if your use case is output-intensive.
Bottom Line
This is a straightforward cost reduction with no reported capability changes. Teams already using Nemotron 3 Ultra should see immediate savings without any integration work. Teams evaluating it now have a slightly easier case to make to finance.
Check the full spec sheet, current pricing, and historical price trend at Nemotron 3 Ultra — live specs & price history.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.