Z.ai Lists GLM 5.3 FlashX: Multimodal Model With 1049k Context at $0.37 Input
What's New
Z.ai added GLM 5.3 FlashX to its lineup on 2026-09-18. It's an open-source multimodal model that accepts text, image, and video inputs and returns text — a combination that's still uncommon at this price tier.
Pricing & Context
At $0.37 per 1M input tokens and $1.25 per 1M output tokens, GLM 5.3 FlashX sits in budget territory. The context window clocks in at 1049k tokens, which is large enough to handle lengthy documents, extended video transcripts, or multi-turn sessions without chunking workarounds.
For a rough sense of workload cost: processing 10M input tokens — think bulk document analysis or a heavy video pipeline — runs about $3.70 in input costs alone. Output-heavy tasks (summarization, generation) will push the bill up faster given the $1.25 output rate, so it's worth estimating your input-to-output ratio before committing.
Who Should Care
- Developers building multimodal pipelines who need video understanding without paying premium-model rates.
- Teams with large context demands — 1049k tokens covers most real-world long-context use cases comfortably.
- Open-source advocates who want model weights access alongside API availability.
If your workload is text-only, cheaper text-specialist models may undercut it. If you need structured outputs or function calling, check whether GLM 5.3 FlashX supports those before switching.
Track It
Pricing in this segment moves quickly. Bookmark GLM 5.3 FlashX — live specs & price history to catch any changes before they affect your budget.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.