Back to Blog

Z.ai Lists GLM 5.3 FlashX: Multimodal Model With 1049k Context at $0.37 Input

EzraSeptember 19, 20261 min read
Z.ai Lists GLM 5.3 FlashX: Multimodal Model With 1049k Context at $0.37 Input

What's New

Z.ai added GLM 5.3 FlashX to its lineup on 2026-09-18. It's an open-source multimodal model that accepts text, image, and video inputs and returns text — a combination that's still uncommon at this price tier.

Pricing & Context

At $0.37 per 1M input tokens and $1.25 per 1M output tokens, GLM 5.3 FlashX sits in budget territory. The context window clocks in at 1049k tokens, which is large enough to handle lengthy documents, extended video transcripts, or multi-turn sessions without chunking workarounds.

For a rough sense of workload cost: processing 10M input tokens — think bulk document analysis or a heavy video pipeline — runs about $3.70 in input costs alone. Output-heavy tasks (summarization, generation) will push the bill up faster given the $1.25 output rate, so it's worth estimating your input-to-output ratio before committing.

Who Should Care

  • Developers building multimodal pipelines who need video understanding without paying premium-model rates.
  • Teams with large context demands — 1049k tokens covers most real-world long-context use cases comfortably.
  • Open-source advocates who want model weights access alongside API availability.

If your workload is text-only, cheaper text-specialist models may undercut it. If you need structured outputs or function calling, check whether GLM 5.3 FlashX supports those before switching.

Track It

Pricing in this segment moves quickly. Bookmark GLM 5.3 FlashX — live specs & price history to catch any changes before they affect your budget.

Ezra, Scout AI Team

E

Ezra

Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.