The Best-Value AI Models Right Now (Data, Not Vibes)
Value is not the cheapest model. It is the most Arena points you buy per dollar of tokens. That distinction is the whole game, and it is why a $0.06 model can beat a $0.15 one on our board while still scoring lower on raw capability.
Below is the current Scout Value Leaderboard: the best value AI models ranked by capability-per-dollar, straight from our pipeline. No vibes, no vendor decks.
Data from our live pipeline, updated September 17, 2026. Prices sync every 6 hours.
How the Value Score works
The method is deliberately boring, because boring is auditable.
Value Score = (Arena Score − 1200) ÷ blended price per 1M tokens, where blended price = 0.75 × input + 0.25 × output.
Two choices matter here. First, we subtract 1200 from the Arena Score so we measure capability above a usable floor, not from zero. A model that barely clears the baseline should not look twice as valuable as it is. Second, the 75/25 input-to-output weighting reflects how most real workloads run: long prompts and context, shorter completions. If your usage skews the other way, weight it yourself; the inputs are all in the table.
Capability comes from LMArena. Prices come from our live market sync, refreshed every 6 hours. Arena Scores © LMArena, licensed CC-BY-4.0, as of 2026-09-02.
Want the full interactive version? It lives on the value leaderboard. To compare specs side by side, start at /models; for raw token costs, /pricing.
The current top 10
| # | Model | Arena | In / Out ($/1M) | Blended | Value Score |
|---|---|---|---|---|---|
| 1 | gpt-oss-120b | 1366 | $0.037 / $0.17 | $0.07 | 2361.6 |
| 2 | Qwen3 30B A3B Instruct 2507 | 1384 | $0.0482 / $0.193 | $0.08 | 2173.5 |
| 3 | DeepSeek V4 Flash 0423 | 1432 | $0.0886 / $0.177 | $0.11 | 2093.9 |
| 4 | Gemma 3 12B | 1334 | $0.05 / $0.15 | $0.08 | 1786.7 |
| 5 | Qwen3.5-Flash | 1398 | $0.065 / $0.26 | $0.11 | 1738.9 |
| 6 | Gemma 4 26B A4B | 1435 | $0.09 / $0.3 | $0.14 | 1648.4 |
| 7 | Gemma 4 31B | 1442 | $0.09 / $0.34 | $0.15 | 1585.6 |
| 8 | gpt-oss-20b | 1287 | $0.03 / $0.13 | $0.06 | 1585.5 |
| 9 | Gemma 3 4B | 1291 | $0.05 / $0.1 | $0.06 | 1449.6 |
| 10 | Gemma 3n 4B | 1305 | $0.06 / $0.12 | $0.08 | 1402.7 |
Two families dominate the board: Google's Gemma line holds six of ten slots, and open-source Qwen and DeepSeek entries take three of the top five. OpenAI's open-weight gpt-oss pair bookends the list at #1 and #8.
Reading the leaderboard
#1 gpt-oss-120b: the value king, for now
At a blended $0.07 and an Arena score of 1366, gpt-oss-120b posts a Value Score of 2361.6 — the highest on the board by a clear margin. It is not the smartest model here; #3, #6 and #7 all score higher on Arena. It wins because its input price of $0.037 is close to rock-bottom while its capability sits comfortably above the pack. That combination is what our score rewards.
The gap to #2 is about 188 points. That is a real cushion, but it is one input-price cut away from tightening.
#2 and #3: capability climbs, price follows
Qwen3 30B A3B Instruct 2507 scores 1384 on Arena — 18 points above the leader — but its blended $0.08 pulls the Value Score to 2173.5. DeepSeek V4 Flash 0423 is the highest-capability model in the top five at 1432, yet its blended $0.11 lands it at 2093.9. The pattern is clean: as you climb the capability ladder, price climbs faster, and value slips.
If you want the smartest model that still counts as good value, DeepSeek V4 Flash is the pick. If you want the best raw dollar efficiency, stay at #1.
The Gemma spread: pick your rung
Google's Gemma family lets you dial capability against cost without leaving one ecosystem:
- Gemma 3 4B and Gemma 3n 4B (#9, #10) — cheapest tier, blended $0.06 and $0.08, Arena in the 1291–1305 band. Fine for classification, extraction and high-volume light tasks.
- Gemma 3 12B (#4) — the mid rung at blended $0.08 and Arena 1334, Value Score 1786.7.
- Gemma 4 26B A4B and Gemma 4 31B (#6, #7) — the capability end, Arena 1435 and 1442, the two smartest models on this entire list. You pay for it: blended $0.14 and $0.15.
Gemma 4 31B has the highest Arena score here (1442) and the lowest Value Score in the top seven. That is the trade in one line.
The floor: gpt-oss-20b and the 4B tier
gpt-oss-20b is the cheapest model on the board at a blended $0.06, but its Arena score of 1287 is also the lowest, which is why it sits at #8 despite the price. Cheap alone does not win here. You need capability above the floor to convert low price into a high Value Score, and the 20B model just clears it.
What each buyer should take from this
Hobbyists and side projects
Start at #1. gpt-oss-120b gives you the best capability-per-dollar on the board, and at a blended $0.07 the bill on a personal project rounds to noise. If you are doing bulk, low-stakes work — tagging, summaries, cleanup — drop to gpt-oss-20b or Gemma 3 4B at a blended $0.06 and save the marginal cents. Browse what is heating up on /trending before you commit.
Startups shipping a product
You care about value and a capability floor high enough that support tickets stay low. That points to DeepSeek V4 Flash (#3, Arena 1432) or Qwen3.5-Flash (#5, Arena 1398). Both clear a strong capability bar while staying at a blended $0.11. If your app is code-heavy, cross-check against our coding view and see how these slot into real stacks before wiring anything into production.
Heavy API users
At scale, the 75/25 input weighting is your friend, because input tokens are usually the bulk of the bill. gpt-oss-120b's $0.037 input is the number to beat. Model the blended price against your input/output ratio, not ours, because a completion-heavy workload can reshuffle this entire order. If a proprietary flagship is still in your stack, the head-to-heads at /vs/deepseek-vs-chatgpt and /vs/claude-vs-chatgpt are worth the five minutes.
The caveat that keeps you honest
This is a live snapshot, not a verdict. Prices sync every 6 hours, and a single cut can reshuffle the top five overnight. A 30% input-price drop on any of #2 through #5 would put real pressure on the leader. Arena scores also move as new models land and voting continues. Treat the leaderboard as a starting shortlist, then confirm against your own traffic mix.
We do not run private benchmarks or claim lab results. Our authority is the pipeline: LMArena for capability, live market sync for price, refreshed on a clock. That is the whole method, and you can re-derive every Value Score above from the table.
FAQ
Why isn't the cheapest model ranked #1?
Because value is capability per dollar, not price alone. gpt-oss-20b is the cheapest at a blended $0.06, but its Arena score of 1287 is the lowest here, so it lands at #8. gpt-oss-120b costs a little more at $0.07 blended and scores 1366, which is why it leads.
How do you calculate the Value Score?
Value Score = (Arena Score − 1200) ÷ blended price per 1M tokens, where blended = 0.75 × input + 0.25 × output. Capability is from LMArena (© LMArena, CC-BY-4.0, as of 2026-09-02); prices are from our live market sync, refreshed every 6 hours.
How often does this ranking change?
Prices update every 6 hours, so the order can shift the same day a provider adjusts token costs. Check the value leaderboard for the current state.
Ezra, Scout AI Team

Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.