Qwen3 Next 80B A3B Thinking API Prices Rise 54% on Inputs and Outputs
What changed
Qwen's Qwen3 Next 80B A3B Thinking model has seen a across-the-board price increase of +54%. Input pricing moves from $0.0975 to $0.15 per 1M tokens; output pricing climbs from $0.78 to $1.20 per 1M tokens.
Who feels it
For low-volume experimentation, the delta is negligible. The pain lands on production workloads that lean heavily on the model's "Thinking" reasoning mode — those tend to generate long output traces, which means output token costs dominate the bill. At $1.20 per 1M output tokens, a pipeline producing 10M output tokens monthly now costs $12,000 on that line item alone, up from $7,800. That's a real budget conversation.
The model's mixture-of-experts architecture (80B total, 3B active) keeps per-call latency and compute lean, which has historically been part of its value proposition. The new pricing narrows that cost advantage against some competitors.
What to do
- Audit output verbosity. Thinking-mode models are chatty by design. Prompt-level instructions to be concise can cut output tokens meaningfully.
- Re-benchmark alternatives. A 54% jump is enough to justify a comparison run against other reasoning-capable models on our site.
- Check your tier. Some API providers apply volume discounts that could offset part of the increase.
For live pricing, current context window specs, and a full price history, see Qwen3 Next 80B A3B Thinking — live specs & price history.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.