GLM 5 API Drops 37%: Input Now $0.60, Output $1.92 per 1M Tokens
What Changed
Z.ai has quietly trimmed GLM 5 API pricing across the board. Input cost moves from $0.95 → $0.60 per 1M tokens, and output from $2.55 → $1.92 per 1M tokens — a -37% cut overall.
Does It Matter for Your Workload?
It depends on your input/output ratio. Workloads that are heavily output-bound (long-form generation, agentic loops, summarisation pipelines) see the softer saving — output dropped ~25%. Teams with input-heavy pipelines (document analysis, retrieval-augmented generation with large contexts) get the sharper win, closer to that full 37% headline figure.
At 100M input tokens a month, you're now saving roughly $35 compared to last week — not life-changing, but real money at the scale where token costs actually show up on a budget spreadsheet.
How It Sits Competitively
At $0.60 input / $1.92 output, GLM 5 is pushing into territory that makes it worth a genuine side-by-side eval if you've been defaulting to pricier alternatives. Whether the quality trade-off holds for your use case is a separate question — benchmarks vary by task.
For a full breakdown of current rates and how they've shifted over time, see GLM 5 — live specs & price history.
Bottom Line
This is a straightforward cost cut, not a capability update. If GLM 5 was already on your shortlist, it just got more defensible in a build-vs-cost conversation. If you ruled it out on price before, worth another look.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.