NVIDIA Nemotron 3 Super API Prices Rise 253%: What Developers Need to Know
The Numbers
NVIDIA has repriced the Nemotron 3 Super API. Input tokens move from $0.085 to $0.30 per million — a +253% increase. Output tokens rise from $0.40 to $0.90 per million. Both changes are live now.
Does This Actually Hurt?
It depends on your usage pattern. Output-heavy workloads (long generations, document drafting, agentic chains) feel the most pressure: at $0.90/M output tokens, a pipeline burning 10M output tokens a month jumps from $4.00 to $9.00 on that line alone. Input-heavy workloads — long-context retrieval, classification, embeddings workarounds — take the bigger proportional hit given the 253% input increase, even though the absolute output rate is still the steeper cost.
At $0.30 input / $0.90 output, Nemotron 3 Super is no longer priced as a budget option. Teams that chose it specifically for cost efficiency should re-run their per-workload math before the next billing cycle.
What To Do
- Audit your token split. If you're output-heavy, the output rate is the bigger lever.
- Compare alternatives. Several models in a similar capability tier sit below these new rates. Use Nemotron 3 Super — live specs & price history to stack it against current competitors side by side.
- Check caching options. If your provider supports prompt caching, high input costs make cache hits more valuable than before.
The model itself hasn't changed — only the price. Whether the performance justifies $0.30/$0.90 depends entirely on what you're getting from it relative to cheaper alternatives now on the board.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.