Nemotron 3 Ultra API Prices Rise: Output Tokens Up 64%
What Changed
NVIDIA has raised API pricing on Nemotron 3 Ultra across both input and output tokens. Input pricing moved from $0.50 to $0.60 per 1M tokens — a modest bump. The bigger hit is on output: costs climbed from $2.20 to $3.60 per 1M tokens, a +64% increase.
Does It Matter for Your Workload?
For most LLM use cases, output tokens dominate total cost. If your application generates long completions — think document drafting, code generation, or agentic chains — this repricing compounds fast. A workload burning 10M output tokens per month just got ~$14 more expensive. At 100M tokens, that's an extra $140/month from this change alone.
Input-heavy workloads (large context retrieval, classification, summarization with short outputs) will feel less pain — the $0.10/1M input increase is relatively contained.
Worth Reconsidering?
At $3.60/1M output tokens, Nemotron 3 Ultra is no longer in budget territory. Before absorbing the increase, it's worth benchmarking whether the model's performance justifies the new rate versus alternatives at similar or lower price points.
Check the full spec sheet and track how this price sits in context: Nemotron 3 Ultra — live specs & price history.
Bottom Line
Price increases this steep on output tokens are worth a deliberate review — not panic, but don't auto-renew assumptions either. Run your actual token split against the new rates before deciding whether to stay or switch.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.