Back to Models
Z.ai: GLM 5.3 Flash
Z AItext9.1 / 10 Overall Rating
GLM-5.3-Flash from Z.ai is a native multimodal model supporting text, image, and video inputs with a 1.3M-token context window. Built on a hybrid sparse-linear attention architecture, it is designed for low-latency coding and agent execution. It offers exceptionally low operational costs at $0.07 per million input tokens.
Context Window
1311K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary
Where it Excels (Pros)
- 1.3M token context window
- Extremely low API pricing
- Supports native video input
- 131K max output tokens
Limitations (Cons)
- Lacks frontier deep reasoning skills
- Linear attention accuracy trade-offs
Benchmark Breakdown
Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%
Pricing Matrix
Input cost / 1M tokens$0.07
Output cost / 1M tokens$0.25
Compare Z.ai: GLM 5.3 Flash with other frontier models: