How is GLM-5.2 Benchmarks constructed?
Published by Deepstory Research as the methodology attached to GLM-5.2 Benchmarks. Scope and source attribution are bounded by that Evidence Brief.
Procedure and scope
Source-backed values for all six charts come from the Z.ai GLM-5.2 launch post (2026-06-16): long-horizon coding (FrontierSWE, PostTrainBench, SWE-Marathon), standard coding (Terminal-Bench 2.1, SWE-bench Pro, NL2Repo, DeepSWE, ProgramBench), reasoning (HLE, CritPt, AIME, HMMT, GPQA-Diamond, IMOAnswerBench), agentic (MCP-Atlas, Tool-Decathlon), the GLM-5.2 vs GLM-5.1 generational leap, and the MTP acceptance-length ablation. Every numeric point carries a sources[].ref and a value_basis naming the table row. Scores are reproduced verbatim from the Full Benchmark Table; cells the vendor published as "-" (not reported) are omitted, never estimated. CAVEAT: these are the model developer’s self-reported figures under their stated harnesses and prompts (see the post’s footnotes); some competitor HLE/CritPt cells are full-set scores. They are comparison-as-published, not an independent eval. Architecture context (not charted): 1M-token context (up from 200K), IndexShare cuts per-token indexer FLOPs ~2.9× at 1M, MIT-licensed open weights. Re-verified 2026-06-22.
Known failure modes and limits
- Independent verification — every figure is vendor self-reported (zai-glm52).
- Real-world reliability, latency or cost (benchmark scores only).