GLM 5.3 Flash
zai-org/GLM-5.3-Flash
Open Source · chat · open-weights
Open
Alert me on changes
Intelligence
Context
1M
Max output
131.1K
Weights
Open
API $/1M
$0.15 / $0.50
Modalities
text · image
Released
25 Aug 2026
Intelligence Index via Artificial Analysis · 0–100, higher is better
License: mit · zai-org/GLM-5.3-Flash
AI summary
● machine-written
Z.ai releases GLM 5.3 Flash, a multimodal mixture-of-experts model
GLM 5.3 Flash is a native multimodal model from Z.ai designed for efficient coding and long-horizon agent tasks. It uses a hybrid sparse and linear attention architecture with 320 billion total parameters but only 18 billion active per token. The model is available as open-weights under MIT license and is priced significantly lower than comparable frontier models.
What's new
- Released under MIT license on Hugging Face for self-hosting
- Mixture-of-experts architecture with 320B total / 18B active parameters
- 1M token context window with 131K max output tokens
- Priced at $0.15 per million input tokens, $0.50 per million output tokens
Best for
Efficient coding tasksLong-horizon agent tasksCost-sensitive token-heavy workloadsLong-context agentic sessions