DeepSeek V4 Flash 0731 Cuts API Prices 36% — Input Now $0.09/1M Tokens
What Changed
DeepSeek has cut API pricing on V4 Flash 0731 by 36% across the board. Input tokens drop from $0.14 to $0.09 per 1M tokens; output tokens fall from $0.28 to $0.18 per 1M tokens.
Who This Actually Affects
The cut matters most for high-throughput workloads — think document processing pipelines, batch classification, or RAG systems generating large output volumes. At $0.18/1M output tokens, the model sits firmly in budget-tier territory, making it worth re-running cost estimates if you benchmarked it before and moved on.
For lighter workloads (occasional queries, low-volume prototyping), the absolute dollar savings will be marginal. The bigger question is whether the quality-per-dollar ratio now beats whatever you're currently running.
Context
Flash-class models are optimized for speed and cost over raw capability. If your use case tolerates that trade-off, a 36% reduction is a meaningful signal to revisit. If you need top-tier reasoning or complex instruction following, this tier probably wasn't your pick regardless of price.
Check current benchmarks, context window, and the full price history before committing budget: DeepSeek V4 Flash 0731 — live specs & price history.
Bottom Line
A -36% cut on both input and output is a real reduction, not a rounding-error adjustment. If DeepSeek V4 Flash 0731 was close to viable for your workload before, run the numbers again — it may have just crossed the threshold.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.