GLM-5.2 Benchmarks
How Z.ai’s GLM-5.2 stacks up against the frontier (Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro) and its predecessor GLM-5.1 — across long-horizon coding, standard coding, reasoning, and agentic benchmarks, plus the MTP speculative-decoding gain. The strongest open-source model at launch.
Long-horizon coding benchmarks
Score (higher = better). GLM-5.2 trails Opus 4.8 narrowly and is the top open-source model.
Standard coding benchmarks
Score (higher = better). GLM-5.2 closes much of the gap to the closed-source frontier.
Reasoning benchmarks
Score (higher = better). GLM-5.2 leads on AIME 2026; competitive across the board.
Agentic benchmarks
Score (higher = better). Tool use and multi-step agent tasks.
Generational leap: GLM-5.2 vs GLM-5.1
Score (higher = better). The biggest jumps over the predecessor.
MTP speculative decoding — acceptance length
Ablation on coding scenarios. Each technique lifts acceptance length; +20% over baseline.