Back to Blog

Qwen3.6 35B A3B Drops Input Price 33% — Output Costs Edge Down Too

EzraOctober 11, 20261 min read
Qwen3.6 35B A3B Drops Input Price 33% — Output Costs Edge Down Too

What Changed

Qwen has repriced Qwen3.6 35B A3B across both token directions. Input drops from $0.15 to $0.10 per 1M tokens — a clean -33% cut. Output moves from $1.00 to $0.95 per 1M tokens, a smaller but still real reduction.

Who Actually Feels This

The input cut matters most for read-heavy workloads: RAG pipelines, document summarisation, long-context classification, or any task where you're stuffing large prompts but generating short responses. A 33% input reduction directly scales with prompt length, so the longer your context, the bigger the saving.

For generation-heavy use cases — chatbots, code completion, long-form drafting — the $0.05 output reduction is modest. On a million output tokens, that's $50 saved, not nothing, but not transformative either.

Quick Maths

A workload burning 10M input tokens and 2M output tokens per month previously cost $1.50 + $2.00 = $3.50. At new rates: $1.00 + $1.90 = $2.90. That's roughly 17% off the combined bill for that ratio — closer to 33% if your workload skews heavily toward input.

Context

This is a competitive segment. MoE models around this parameter count are under sustained pricing pressure. Qwen trimming here is a signal they're defending volume, not margin. If you're already using this model, there's nothing to migrate — the saving is automatic via API.

Check Qwen3.6 35B A3B — live specs & price history to compare against alternatives on the same task types before assuming this is the best fit for your budget.

Ezra, Scout AI Team

E

Ezra

Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.