Qwen3 VL 235B A22B Thinking API Pricing: Input Cost Jumps 145%
What Changed
Qwen has repriced its Qwen3 VL 235B A22B Thinking — live specs & price history. Input tokens climbed from $0.40 to $0.98 per 1M tokens — a +145% increase. Output pricing moved in the opposite direction, dropping marginally from $4.00 to $3.95 per 1M tokens.
The net effect depends almost entirely on your input-to-output ratio.
Does This Matter for Your Workload?
For most vision-language tasks — document parsing, image captioning, visual QA — prompts tend to be input-heavy, especially once you factor in image token counts. At $0.98/1M input tokens, a workload that was costing you $10/day on inputs alone now costs roughly $24.50. That's not trivial at scale.
The slight output price cut ($4.00 → $3.95) offers negligible relief. Unless your pipeline is unusually output-dense, this repricing is a net cost increase for the vast majority of users.
What to Do Next
- Audit your token split. If your input tokens consistently outnumber output tokens 3:1 or more, the effective per-request cost jump will be significant.
- Check alternatives. Our comparison tools let you filter multimodal models by price tier. If you're not specifically relying on the 235B parameter scale or the thinking-mode reasoning, smaller MoE variants may now offer a better cost-to-capability ratio.
- Watch for further moves. A +145% input reprice in a single step is a notable adjustment. It's worth monitoring whether this stabilizes or signals a broader pricing reset for the Qwen3 VL family.
The output token economy remains competitive at $3.95/1M, but input costs are now the deciding factor for budget-conscious teams.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.