Skip to content

GLM 5.3 Flash BF16

zai-org/GLM-5.3-Flash-BF16
Open Source · chat · open-weights
Open Alert me on changes
Intelligence
Context
1.3M
Max output
131.1K
Weights
Open
API $/1M
$0.15 / $0.50
Modalities
text · image
Released
25 Aug 2026
Intelligence Index via Artificial Analysis · 0–100, higher is better
License: mit · zai-org/GLM-5.3-Flash-BF16
Download image Share on X Share on LinkedIn
AI summary
● machine-written

Z.ai releases GLM-5.3-Flash, sparse MoE model with 1M context window

GLM-5.3-Flash is a native multimodal model from Z.ai featuring sparse mixture-of-experts architecture with 320B total parameters and 18B active per token. The model supports text and image input, includes always-on reasoning capabilities, and is priced at $0.15 per million input tokens and $0.50 per million output tokens. It is suited for efficient coding and long-horizon agent tasks with a 1M-token context window.

What's new
  • Released August 26, 2026 under MIT license on Hugging Face
  • Achieves 63.4 Pass@1 on DeepSWE v1.1, matching independently verified benchmarks
  • Hybrid sparse and linear attention architecture reduces compute overhead
  • Promotional pricing at 50% discount through September 9, 2026
  • Available across 20+ providers via OpenRouter with load balancing
Best for
Efficient coding and code generation tasksLong-horizon agent tasks with tool callingAgentic workflows requiring automation and reasoningLong-context applications with 1M token window
Sources

Source: https://huggingface.co/zai-org/GLM-5.3-Flash-BF16