---
title: "GLM-5.2 Benchmarks"
slug: glm-5-2-benchmarks
type: data-index
sector: ai-models
canonical_url: https://deepstoryresearch.com/data/glm-5-2-benchmarks
series: [coding_long, coding_std, agentic, reasoning, generational, mtp_ablation]
series_count: 6
data_point_count: 28
data_as_of: 2026-06
source_quality: curated_snapshot
tier: free
generated_at: 2026-08-11
---

# GLM-5.2 Benchmarks
_GLM-5.2 vs the frontier across coding, reasoning and agentic benchmarks_

> Deepstory Research context file · **Free tier** · <https://deepstoryresearch.com/data/glm-5-2-benchmarks>
> Self-contained AI briefing. Drop into an LLM window or RAG pipeline.

## AI research use contract

- Treat this file as source-grounded context, not a live database. Quote `source_ref` and `data_as_of` when using numbers.
- Separate `FACT`, `ESTIMATE`, `FORECAST`, `TARGET`, and `INFERENCE`. If a row basis is unclear, say it is unclear.
- Prefer T1/T2 sources for hard facts; use T3/T4 sources as context that should be checked before money, legal, medical, operational, or policy decisions.
- Keep answer structure clean: first say what the data says, then what you infer, then what would change the conclusion.

## Research upgrade checklist

- Add independent replication status for each benchmark, especially where the only source is the model lab.
- Separate coding, agentic, reasoning, long-context, and cost-efficiency claims.
- Include benchmark definitions and score direction so AI tools do not mix pass@k, percentile, and composite scores.
- Add a China/export-control note where infrastructure or access constraints matter to users.

## TL;DR — index summary

Z.ai's GLM-5.2 closes most of the gap to the Western frontier: on FrontierSWE it scores 74.4 vs Claude Opus 4.8's 75.1 and GPT-5.5's 72.6 (and Gemini's 39.6), and it is a generational leap over GLM-5.1 (SWE-Marathon 13.0 vs 1.0). Read every number with one caveat: all of it comes from a single vendor source (Z.ai's own release), so it is a self-reported benchmark table, not independent evaluation.

## What this index covers

Six series comparing GLM-5.2 to Opus 4.8, GPT-5.5, Gemini and its own predecessor GLM-5.1: long-horizon coding, standard coding, reasoning, agentic, a generational 5.2-vs-5.1 delta, and a multi-token-prediction (MTP) ablation.

**Series in this index:**

- `coding_long` — FrontierSWE-class long-task coding scores across models.
- `coding_std` — Terminal-Bench / standard coding benchmarks.
- `agentic` — Agentic/tool-use benchmarks (e.g. MCP-Atlas).
- `reasoning` — Reasoning benchmarks (e.g. HLE).
- `generational` — GLM-5.2 vs GLM-5.1 on the same tasks.
- `mtp_ablation` — Effect of multi-token prediction on accepted length.

**Entities tracked:** Z.ai (GLM), Anthropic (Claude Opus 4.8), OpenAI (GPT-5.5), Google (Gemini).

## Key findings

- FrontierSWE: GLM-5.2 74.4 vs Opus 4.8 75.1 vs GPT-5.5 72.6 vs Gemini 39.6 vs GLM-5.1 30.5 (zai-glm52).
- Terminal-Bench 2.1: GLM-5.2 81 vs Opus 4.8 85 vs GPT-5.5 84 vs Gemini 74 (zai-glm52).
- Generational leap: SWE-Marathon 13.0 (5.2) vs 1.0 (5.1) (zai-glm52).
- MTP ablation quantifies the multi-token-prediction contribution to accepted length (zai-glm52).

## Evidence basis map

The free tables above show `source_ref` but not the row-level `value_basis`. This section mirrors the visible rows only, so AI tools can separate observed values from estimates, forecasts, targets, and derived values.

| series | row | source_ref | value_basis |
| --- | --- | --- | --- |
| coding_long | benchmark=FrontierSWE; glm52=74.4; glm51=30.5; opus48=75.1; gpt55=72.6; gemini=39.6 | zai-glm52 | Full Benchmark Table — FrontierSWE (dominance as of 2026-06-16) |
| coding_long | benchmark=PostTrainBench; glm52=34.3; glm51=20.1; opus48=37.2; gpt55=28.4; gemini=21.6 | zai-glm52 | Full Benchmark Table — PostTrainBench |
| coding_long | benchmark=SWE-Marathon; glm52=13; glm51=1; opus48=26; gpt55=12; gemini=4 | zai-glm52 | Full Benchmark Table — SWE-Marathon |
| coding_std | benchmark=Terminal-Bench 2.1; glm52=81; glm51=63.5; opus48=85; gpt55=84; gemini=74 | zai-glm52 | Full Benchmark Table — Terminal-Bench 2.1 (Terminus-2) |
| coding_std | benchmark=SWE-bench Pro; glm52=62.1; glm51=58.4; opus48=69.2; gpt55=58.6; gemini=54.2 | zai-glm52 | Full Benchmark Table — SWE-bench Pro |
| coding_std | benchmark=NL2Repo; glm52=48.9; glm51=42.7; opus48=69.7; gpt55=50.7; gemini=33.4 | zai-glm52 | Full Benchmark Table — NL2Repo |
| coding_std | benchmark=DeepSWE; glm52=46.2; glm51=18; opus48=58; gpt55=70; gemini=10 | zai-glm52 | Full Benchmark Table — DeepSWE |
| coding_std | benchmark=ProgramBench; glm52=63.7; glm51=50.9; opus48=71.9; gpt55=70.8; gemini=39.5 | zai-glm52 | Full Benchmark Table — ProgramBench |
| agentic | benchmark=MCP-Atlas; glm52=76.8; glm51=71.8; opus48=77.8; gpt55=75.3; gemini=69.2 | zai-glm52 | Full Benchmark Table — MCP-Atlas (public set) |
| agentic | benchmark=Tool-Decathlon; glm52=48.2; glm51=40.7; opus48=59.9; gpt55=55.6; gemini=48.8 | zai-glm52 | Full Benchmark Table — Tool-Decathlon |
| reasoning | benchmark=HLE; glm52=40.5; glm51=31; opus48=49.8; gpt55=41.4; gemini=45 | zai-glm52 | Full Benchmark Table — HLE (Opus/GPT marked full-set) |
| reasoning | benchmark=HLE w/ Tools; glm52=54.7; glm51=52.3; opus48=57.9; gpt55=52.2; gemini=51.4 | zai-glm52 | Full Benchmark Table — HLE with tools |
| reasoning | benchmark=CritPt; glm52=20.9; glm51=4.6; opus48=20.9; gpt55=27.1; gemini=17.7 | zai-glm52 | Full Benchmark Table — CritPt |
| reasoning | benchmark=AIME 2026; glm52=99.2; glm51=95.3; opus48=95.7; gpt55=98.3; gemini=98.2 | zai-glm52 | Full Benchmark Table — AIME 2026 |
| reasoning | benchmark=HMMT Feb 2026; glm52=92.5; glm51=82.6; opus48=96.7; gpt55=96.7; gemini=87.3 | zai-glm52 | Full Benchmark Table — HMMT Feb 2026 |
| generational | benchmark=SWE-Marathon; glm52=13; glm51=1 | zai-glm52 | 5.2 vs 5.1: 13.0 vs 1.0 |
| generational | benchmark=CritPt; glm52=20.9; glm51=4.6 | zai-glm52 | 5.2 vs 5.1: 20.9 vs 4.6 |
| generational | benchmark=DeepSWE; glm52=46.2; glm51=18 | zai-glm52 | 5.2 vs 5.1: 46.2 vs 18.0 |
| generational | benchmark=FrontierSWE; glm52=74.4; glm51=30.5 | zai-glm52 | 5.2 vs 5.1: 74.4 vs 30.5 |
| generational | benchmark=PostTrainBench; glm52=34.3; glm51=20.1 | zai-glm52 | 5.2 vs 5.1: 34.3 vs 20.1 |
| mtp_ablation | step=Baseline; acc_len=4.56 | zai-glm52 | MTP ablation table — baseline 4.56 |
| mtp_ablation | step=+ IndexShare + KVShare; acc_len=5.1 | zai-glm52 | MTP ablation table — 5.10 |
| mtp_ablation | step=+ Rejection Sampling; acc_len=5.29 | zai-glm52 | MTP ablation table — 5.29 |
| mtp_ablation | step=+ End-to-end TV Loss; acc_len=5.47 | zai-glm52 | MTP ablation table — 5.47 (+20% over baseline) |

## The series (per-chart briefings)

### Long-horizon coding — grouped bar
**Direct answer:** FrontierSWE-class long-task coding scores across models.
**Time bracket:** snapshot

| benchmark | glm52 | glm51 | opus48 | gpt55 | gemini | source_ref |
| --- | --- | --- | --- | --- | --- | --- |
| FrontierSWE | 74.4 | 30.5 | 75.1 | 72.6 | 39.6 | zai-glm52 |
| PostTrainBench | 34.3 | 20.1 | 37.2 | 28.4 | 21.6 | zai-glm52 |
| SWE-Marathon | 13 | 1 | 26 | 12 | 4 | zai-glm52 |

**Read:** GLM-5.2 is within ~1 point of Opus 4.8 and ahead of GPT-5.5.

### Standard coding — grouped bar
**Direct answer:** Terminal-Bench / standard coding benchmarks.
**Time bracket:** snapshot

| benchmark | glm52 | glm51 | opus48 | gpt55 | gemini | source_ref |
| --- | --- | --- | --- | --- | --- | --- |
| Terminal-Bench 2.1 | 81 | 63.5 | 85 | 84 | 74 | zai-glm52 |
| SWE-bench Pro | 62.1 | 58.4 | 69.2 | 58.6 | 54.2 | zai-glm52 |
| NL2Repo | 48.9 | 42.7 | 69.7 | 50.7 | 33.4 | zai-glm52 |
| DeepSWE | 46.2 | 18 | 58 | 70 | 10 | zai-glm52 |
| ProgramBench | 63.7 | 50.9 | 71.9 | 70.8 | 39.5 | zai-glm52 |

**Read:** A few points behind Opus/GPT, well ahead of Gemini.

### Agentic — grouped bar
**Direct answer:** Agentic/tool-use benchmarks (e.g. MCP-Atlas).
**Time bracket:** snapshot

| benchmark | glm52 | glm51 | opus48 | gpt55 | gemini | source_ref |
| --- | --- | --- | --- | --- | --- | --- |
| MCP-Atlas | 76.8 | 71.8 | 77.8 | 75.3 | 69.2 | zai-glm52 |
| Tool-Decathlon | 48.2 | 40.7 | 59.9 | 55.6 | 48.8 | zai-glm52 |

**Read:** Near-parity on agentic tasks.

### Reasoning — grouped bar
**Direct answer:** Reasoning benchmarks (e.g. HLE).
**Time bracket:** snapshot

| benchmark | glm52 | glm51 | opus48 | gpt55 | gemini | source_ref |
| --- | --- | --- | --- | --- | --- | --- |
| HLE | 40.5 | 31 | 49.8 | 41.4 | 45 | zai-glm52 |
| HLE w/ Tools | 54.7 | 52.3 | 57.9 | 52.2 | 51.4 | zai-glm52 |
| CritPt | 20.9 | 4.6 | 20.9 | 27.1 | 17.7 | zai-glm52 |
| AIME 2026 | 99.2 | 95.3 | 95.7 | 98.3 | 98.2 | zai-glm52 |
| HMMT Feb 2026 | 92.5 | 82.6 | 96.7 | 96.7 | 87.3 | zai-glm52 |

_+2 more rows on the live page and in the full working dataset._

**Read:** The widest remaining gap to Opus/Gemini.
**Caveat:** Some competitor scores are marked full-set vs public-set — not always like-for-like.

### Generational delta — bar
**Direct answer:** GLM-5.2 vs GLM-5.1 on the same tasks.
**Time bracket:** snapshot

| benchmark | glm52 | glm51 | source_ref |
| --- | --- | --- | --- |
| SWE-Marathon | 13 | 1 | zai-glm52 |
| CritPt | 20.9 | 4.6 | zai-glm52 |
| DeepSWE | 46.2 | 18 | zai-glm52 |
| FrontierSWE | 74.4 | 30.5 | zai-glm52 |
| PostTrainBench | 34.3 | 20.1 | zai-glm52 |

_+2 more rows on the live page and in the full working dataset._

**Read:** A large jump generation-on-generation.

### MTP ablation — bar
**Direct answer:** Effect of multi-token prediction on accepted length.
**Time bracket:** ablation

| step | acc_len | source_ref |
| --- | --- | --- |
| Baseline | 4.56 | zai-glm52 |
| + IndexShare + KVShare | 5.1 | zai-glm52 |
| + Rejection Sampling | 5.29 | zai-glm52 |
| + End-to-end TV Loss | 5.47 | zai-glm52 |

**Read:** Shows where the speed/quality gains come from architecturally.

## Cross-series synthesis — why this matters

The coding and agentic series together say the open-challenger gap to the Western frontier has narrowed to low single digits, while the reasoning series says the last gap is in hard reasoning. The generational series (5.2 vs 5.1) shows the jump was fast — but because all six series share one vendor source, the right read is "impressive self-reported parity pending independent confirmation," not "confirmed frontier."

## Entities & relationships

| Entity | What they do | Key stat | Relationship |
| --- | --- | --- | --- |
| Z.ai (GLM) | Chinese open model lab | GLM-5.2 near-frontier (self-reported) | Challenger to Anthropic/OpenAI/Google |
| Anthropic — Opus 4.8 | Frontier comparator | Leads most benchmarks here | The bar GLM-5.2 measures against |
| OpenAI — GPT-5.5 | Frontier comparator | Neck-and-neck with GLM-5.2 on coding | Second reference point |
| Google — Gemini | Frontier comparator | Trails on the coding set here | Third reference point |

## Timeline of key events & decisions

- **2026-06-16** — GLM-5.2 benchmark table published → the single dated source for every number in this index

## Cross-industry ripple

- **Open-weight AI:** A near-frontier open challenger pressures Western labs' pricing and moat (see frontier-ai-race). _(inference)_
- **Geopolitics / export policy:** Chinese frontier-parity claims feed the compute-export and AI-competition debate. _(inference)_

## Non-obvious reads (interpretation)

_This section is interpretation, not sourced fact — each item names the observation it is built on and a confidence level._

- **Observation:** All six series trace to one vendor source (zai-glm52).
  **Read:** Treat the parity claim as a hypothesis to be checked against independent evals (LMArena, Artificial Analysis) — vendor benchmark tables systematically favour the vendor. _(confidence: high)_

## Glossary / key terms

- **GLM-5.2** — Z.ai's 2026 frontier-class model.
- **FrontierSWE / SWE-Marathon** — Long-horizon software-engineering benchmarks.
- **Terminal-Bench** — A terminal/agentic coding benchmark.
- **HLE** — "Humanity's Last Exam" — a hard reasoning benchmark.
- **MTP** — Multi-token prediction — an architecture technique for speed/quality.

## How to use this with AI

Paste this file into an LLM context window (or a RAG store) and ask cross-series questions. The tables carry a `source_ref` per row; the source registry below maps each ref to a named source and a trust tier.

### Suggested prompts (multi-series)

```text
Using all six series, write a balanced verdict on whether GLM-5.2 is "at the frontier," foregrounding the single-source caveat.
```

```text
Compare GLM-5.2's coding vs reasoning gaps to Opus 4.8 and say where the model is genuinely competitive.
```

```text
From the generational series, characterise how far GLM moved from 5.1 to 5.2.
```

## Sources & trust

| ref | source | trust tier | url |
| --- | --- | --- | --- |
| zai-glm52 | Z.ai — GLM-5.2: Built for Long-Horizon Tasks (2026-06-16) | T2 · Primary filing / official body | https://z.ai/blog/glm-5.2 |

## Caveats & what this index cannot answer

- Independent verification — every figure is vendor self-reported (zai-glm52).
- Real-world reliability, latency or cost (benchmark scores only).
- Estimates and forward targets are labelled in the working dataset; never read a labelled estimate or forecast as a settled figure.
- This is an observational data index, not investment, legal, or medical advice.

## Data freshness & methodology

- **Last updated:** 2026-08-11
- **Data as of:** 2026-06
- **What changed most recently:** GLM-5.2 table (2026-06-16) shows near-parity on coding/agentic; reasoning still trails.

**Methodology (as seeded):**

> Source-backed values for all six charts come from the Z.ai GLM-5.2 launch post
> (2026-06-16): long-horizon coding (FrontierSWE, PostTrainBench, SWE-Marathon),
> standard coding (Terminal-Bench 2.1, SWE-bench Pro, NL2Repo, DeepSWE,
> ProgramBench), reasoning (HLE, CritPt, AIME, HMMT, GPQA-Diamond, IMOAnswerBench),
> agentic (MCP-Atlas, Tool-Decathlon), the GLM-5.2 vs GLM-5.1 generational leap, and
> the MTP acceptance-length ablation.
> 
> Every numeric point carries a sources[].ref and a value_basis naming the table row.
> Scores are reproduced verbatim from the Full Benchmark Table; cells the vendor
> published as "-" (not reported) are omitted, never estimated.
> 
> CAVEAT: these are the model developer’s self-reported figures under their stated
> harnesses and prompts (see the post’s footnotes); some competitor HLE/CritPt cells
> are full-set scores. They are comparison-as-published, not an independent eval.
> Architecture context (not charted): 1M-token context (up from 200K), IndexShare
> cuts per-token indexer FLOPs ~2.9× at 1M, MIT-licensed open weights.
> Re-verified 2026-06-22.

---

_Free tier. The full row-level dataset and per-source detail live on the live page and in the working file. Canonical: <https://deepstoryresearch.com/data/glm-5-2-benchmarks>. Deepstory Research · https://deepstoryresearch.com_
