Back to Blog

Qwen3 VL 235B A22B Thinking API Pricing: Input Cost Jumps 145%

EzraAugust 3, 20261 min read
Qwen3 VL 235B A22B Thinking API Pricing: Input Cost Jumps 145%

What Changed

Qwen has repriced its Qwen3 VL 235B A22B Thinking — live specs & price history. Input tokens climbed from $0.40 to $0.98 per 1M tokens — a +145% increase. Output pricing moved in the opposite direction, dropping marginally from $4.00 to $3.95 per 1M tokens.

The net effect depends almost entirely on your input-to-output ratio.

Does This Matter for Your Workload?

For most vision-language tasks — document parsing, image captioning, visual QA — prompts tend to be input-heavy, especially once you factor in image token counts. At $0.98/1M input tokens, a workload that was costing you $10/day on inputs alone now costs roughly $24.50. That's not trivial at scale.

The slight output price cut ($4.00 → $3.95) offers negligible relief. Unless your pipeline is unusually output-dense, this repricing is a net cost increase for the vast majority of users.

What to Do Next

  • Audit your token split. If your input tokens consistently outnumber output tokens 3:1 or more, the effective per-request cost jump will be significant.
  • Check alternatives. Our comparison tools let you filter multimodal models by price tier. If you're not specifically relying on the 235B parameter scale or the thinking-mode reasoning, smaller MoE variants may now offer a better cost-to-capability ratio.
  • Watch for further moves. A +145% input reprice in a single step is a notable adjustment. It's worth monitoring whether this stabilizes or signals a broader pricing reset for the Qwen3 VL family.

The output token economy remains competitive at $3.95/1M, but input costs are now the deciding factor for budget-conscious teams.

Ezra, Scout AI Team

E

Ezra

Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.

Qwen3 VL 235B A22B Thinking API Pricing: Input Cost Jumps 145% | AIToolScout