Google Cuts Gemini 3.6 Flash API Prices by 50% — What It Means for Your Bill
The Numbers
Google has halved Gemini 3.6 Flash's API pricing across the board. Input tokens drop from $1.50 to $0.75 per 1M tokens; output tokens drop from $7.50 to $3.75 per 1M tokens — a clean -50% on both sides.
No capability changes, no context window adjustments announced alongside it. This is a straight cost reduction.
Who Actually Benefits
The output price is where most production workloads feel the pain, so the $3.75 new rate is the more meaningful number. If you're running a summarization pipeline, a chat assistant, or any task with verbose responses, you're looking at roughly half the previous output spend — automatically, no code changes required.
For read-heavy or classification workloads (short outputs, lots of input), the $0.75 input rate matters more. Either way, the 50% cut is uniform, so the math is simple: whatever you spent last month, budget half.
Worth Comparing
A 50% cut repositions Flash more competitively against other fast, affordable models in the mid-tier segment. Before committing budget, it's worth checking whether output quality and latency still hold up for your specific use case — pricing is only one variable.
See current rates, context limits, and how they've shifted over time on the Gemini 3.6 Flash — live specs & price history page.
Bottom Line
If you're already using Gemini 3.6 Flash, this is a passive win — your costs drop without any action. If you've been on the fence, the new output rate of $3.75 per 1M tokens makes the model noticeably cheaper to evaluate in a real workload context.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.