Qwen3.6 35B A3B Drops Input Price 33% — Output Costs Edge Down Too
What Changed
Qwen has repriced Qwen3.6 35B A3B across both token directions. Input drops from $0.15 to $0.10 per 1M tokens — a clean -33% cut. Output moves from $1.00 to $0.95 per 1M tokens, a smaller but still real reduction.
Who Actually Feels This
The input cut matters most for read-heavy workloads: RAG pipelines, document summarisation, long-context classification, or any task where you're stuffing large prompts but generating short responses. A 33% input reduction directly scales with prompt length, so the longer your context, the bigger the saving.
For generation-heavy use cases — chatbots, code completion, long-form drafting — the $0.05 output reduction is modest. On a million output tokens, that's $50 saved, not nothing, but not transformative either.
Quick Maths
A workload burning 10M input tokens and 2M output tokens per month previously cost $1.50 + $2.00 = $3.50. At new rates: $1.00 + $1.90 = $2.90. That's roughly 17% off the combined bill for that ratio — closer to 33% if your workload skews heavily toward input.
Context
This is a competitive segment. MoE models around this parameter count are under sustained pricing pressure. Qwen trimming here is a signal they're defending volume, not margin. If you're already using this model, there's nothing to migrate — the saving is automatic via API.
Check Qwen3.6 35B A3B — live specs & price history to compare against alternatives on the same task types before assuming this is the best fit for your budget.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.