Skip to content

GLM 5.3 Flash

zai-org/GLM-5.3-Flash
Open Source · chat · open-weights
Open Alert me on changes
Intelligence
Context
1M
Max output
131.1K
Weights
Open
API $/1M
$0.15 / $0.50
Modalities
text · image
Released
25 Aug 2026
Intelligence Index via Artificial Analysis · 0–100, higher is better
License: mit · zai-org/GLM-5.3-Flash
Download image Share on X Share on LinkedIn
AI summary
● machine-written

Z.ai releases GLM 5.3 Flash, a multimodal mixture-of-experts model

GLM 5.3 Flash is a native multimodal model from Z.ai designed for efficient coding and long-horizon agent tasks. It uses a hybrid sparse and linear attention architecture with 320 billion total parameters but only 18 billion active per token. The model is available as open-weights under MIT license and is priced significantly lower than comparable frontier models.

What's new
  • Released under MIT license on Hugging Face for self-hosting
  • Mixture-of-experts architecture with 320B total / 18B active parameters
  • 1M token context window with 131K max output tokens
  • Priced at $0.15 per million input tokens, $0.50 per million output tokens
Best for
Efficient coding tasksLong-horizon agent tasksCost-sensitive token-heavy workloadsLong-context agentic sessions
Sources

Source: https://huggingface.co/zai-org/GLM-5.3-Flash