DeepSeek V4 Flash 0731 Raises Output Price 50% — What It Costs You Now
The Price Change
DeepSeek has updated API pricing for V4 Flash 0731. The output token rate rose from $0.12 to $0.18 per 1M tokens — a clean +50% increase. Input pricing moved more modestly, from $0.06 to $0.065 per 1M tokens.
See the full details and track future changes on the DeepSeek V4 Flash 0731 — live specs & price history.
Does It Actually Matter for Your Workload?
It depends heavily on your input/output ratio. Most LLM workloads are output-heavy — completions, summaries, code generation — which means the 50% output hike will hit harder than the modest input adjustment.
A rough illustration: if you were spending $1.20 in output costs per 1B output tokens before, you're now spending $1.80. For low-volume or input-heavy use cases (classification, embeddings-adjacent tasks), the real-world delta is smaller.
Should You Switch?
V4 Flash 0731 was attractive partly because of its aggressive pricing. At $0.18/1M output tokens it's still on the cheaper end of the market, but the gap to competitors has narrowed. If your pipeline is output-intensive and cost-sensitive, it's worth running a quick comparison against other models in a similar capability tier — our model comparison tool lets you filter by price and benchmark scores side by side.
If you're locked into V4 Flash 0731 for quality or latency reasons, the new pricing is unlikely to be a dealbreaker at moderate scale. At high scale (hundreds of millions of tokens/month), re-evaluating your model mix makes sense.
Bottom Line
A 50% output price hike is not trivial on paper, but V4 Flash 0731 remains competitively priced in absolute terms. Audit your token ratios, run the numbers for your actual volume, and check alternatives before assuming you need to act.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.