Back to Models

Z.ai: GLM 5.3 Flash

Z AItext
9.1 / 10 Overall Rating

GLM-5.3-Flash from Z.ai is a native multimodal model supporting text, image, and video inputs with a 1.3M-token context window. Built on a hybrid sparse-linear attention architecture, it is designed for low-latency coding and agent execution. It offers exceptionally low operational costs at $0.07 per million input tokens.

Context Window
1311K
Knowledge Cutoff
2025
Max Output
16K
License
Proprietary

Where it Excels (Pros)

  • 1.3M token context window
  • Extremely low API pricing
  • Supports native video input
  • 131K max output tokens

Limitations (Cons)

  • Lacks frontier deep reasoning skills
  • Linear attention accuracy trade-offs

Benchmark Breakdown

Reasoning (MMLU)82%
Coding (HumanEval)78%
Mathematics (MATH)76%
Long-Context Retrieval95%

Pricing Matrix

Input cost / 1M tokens$0.07
Output cost / 1M tokens$0.25

Compare Z.ai: GLM 5.3 Flash with other frontier models:

Related Skills, Tools & Automations