DeepSeek V4 Flash API Prices Rise 49%: What It Costs You Now
The Numbers
DeepSeek has quietly pushed V4 Flash API prices up +49% across the board:
- Input: $0.0938 → $0.14 per 1M tokens
- Output: $0.188 → $0.28 per 1M tokens
For context, see DeepSeek V4 Flash — live specs & price history.
Does This Actually Hurt?
It depends on your volume. At the new rates, running 10M input tokens a month jumps from roughly $0.94 to $1.40 — a $0.46 difference. Multiply that by heavier workloads (100M+ tokens) and the delta starts to sting, especially for high-throughput pipelines where output costs dominate.
Output pricing is the one to watch. At $0.28 per 1M tokens, verbose-response workloads — summarization, code generation, long-form drafting — will feel this increase more than classification or embedding-adjacent tasks.
What To Do
- Low-volume users: Probably fine to stay put. The absolute dollar change is small below ~5M tokens/month.
- Mid-to-high-volume users: Worth benchmarking alternatives. Several models in the same speed/capability tier remain below the old V4 Flash price on output — check our comparison filters.
- Production pipelines: Reprice your cost-per-query estimates now, especially if you negotiated budgets against the $0.188 output rate.
The Bigger Picture
DeepSeek had been aggressive on pricing since launch. A 49% increase signals the honeymoon period for rock-bottom Flash rates may be ending. Whether quality or capacity constraints drove this isn't confirmed — but the direction is clear.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.