Z.ai GLM 5.3 Flash Listed: $0.075 Input, 1311k Context, Video-In Support
What It Is
Z.ai quietly listed GLM 5.3 Flash on 2026-08-26. It's an open-source multimodal model that accepts text, image, and video inputs and returns text — a combination that's still uncommon at this price point.
Pricing
- Input: $0.075 per 1M tokens
- Output: $0.25 per 1M tokens
That output price sits comfortably in "cheap Flash-class" territory. For comparison, if you're running a workload that generates 10M output tokens a month, you're looking at $2.50 — low enough that cost probably isn't your primary filter here.
Context Window
The 1311k-token context is the headline number worth paying attention to. That's large enough to feed in entire codebases, lengthy legal documents, or extended video transcripts in a single call — use cases where you'd otherwise need chunking logic or a more expensive model.
Who Should Care
Developers working on document-heavy pipelines or video understanding will find the input modalities and context depth genuinely useful. The open-source status also means you can self-host if the API pricing or data-residency terms don't suit you.
Casual users and teams doing straightforward text generation probably don't need 1311k context — standard alternatives with smaller windows will serve fine and may have stronger benchmark track records.
Caveats
Listing date is recent, so real-world throughput, rate limits, and quality benchmarks aren't yet widely documented. Treat this as one to test against your specific workload rather than a proven production default.
Check GLM 5.3 Flash — live specs & price history for up-to-date pricing and any spec changes as the model matures.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.