This is the complete set of published GLM benchmarks in one place: 17 evaluations across 8 models, covering reasoning, coding and agentic tool use, plus the three long-horizon evaluations that were run by third parties. Everything on this page is a result as published by the model authors or by the organisation that ran the evaluation. glmmodel.com reports these figures; it does not run them.
That distinction is not a disclaimer for its own sake. Benchmark scores are harness-dependent to a degree that surprises people — the same model on the same evaluation moves several points between Terminus-2 and Claude Code — so the methodology section at the bottom of this page is arguably more useful than the table itself. Read the numbers as reported capability under stated conditions, not as a promise about your workload.
All figures on this page are results as published by the model authors. glmmodel.com reports them; it does not run these evaluations.
Dominance, tasks up to 20 hours
Top open-source modelUp to 10 hours on a single H100
2nd overallUltra-long-horizon SWE, up to 10 hours
2nd to the Opus seriesThe long-horizon trio is what the current GLM model generation was built for, and all three were run by external organisations rather than the model authors. FrontierSWE measures dominance on tasks running up to twenty hours; PostTrainBench runs up to ten hours on a single H100; SWE-Marathon is ultra-long-horizon software engineering, also up to ten hours. GLM-5.2 is the highest-ranked open-source model on all three. It trails Claude Opus 4.8 by seven tenths of a point on FrontierSWE, beats GPT-5.5 there by 1.8, and beats the previous-generation Opus 4.7 by more than eleven.
Comparing GLM-5.2 against GLM-5.1 isolates what one generation of work produced, and comparing both against the best closed model on each row shows how much distance is left. The delta chips below are the published GLM-5.2 score minus the published GLM-5.1 score. Note how uneven they are: 3.7 points on SWE-bench Pro against 28.2 points on DeepSWE. Evaluations that reward sustained execution moved far more than evaluations that reward a single good patch.
The full table below is the authoritative version, grouped into reasoning, coding and agentic sections. The GLM-5.2 column is highlighted. A dash means the model authors published no score for that model on that benchmark, and this site does not fill those cells with estimates. Scroll horizontally to see every model column.
| Benchmark | GLM-5.2 | GLM-5.1 | Qwen3.7-Max | MiniMax M3 | DeepSeek-V4-Pro | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|---|---|
| Reasoning | ||||||||
| HLE | 40.5 | 31.0 | 41.4 | 37.0 | 37.7 | 49.8* | 41.4* | 45.0 |
| HLE w/ Tools | 54.7 | 52.3 | 53.5 | – | 48.2 | 57.9* | 52.2* | 51.4* |
| CritPt | 20.9 | 4.6 | 13.4 | 3.7 | 12.9 | 20.9 | 27.1 | 17.7 |
| AIME 2026 | 99.2 | 95.3 | 97.0 | – | 94.6 | 95.7 | 98.3 | 98.2 |
| HMMT Nov. 2025 | 94.4 | 94.0 | 95.0 | 84.4 | 94.4 | 96.5 | 96.5 | 94.8 |
| HMMT Feb. 2026 | 92.5 | 82.6 | 97.1 | 84.4 | 95.2 | 96.7 | 96.7 | 87.3 |
| IMOAnswerBench | 91.0 | 83.8 | 90.0 | – | 89.8 | 83.5 | – | 81.0 |
| GPQA-Diamond | 91.2 | 86.2 | 90.0 | 93.0 | 90.1 | 93.6 | 93.6 | 94.3 |
| Coding | ||||||||
| SWE-bench Pro | 62.1 | 58.4 | 60.6 | 59.0 | 55.4 | 69.2 | 58.6 | 54.2 |
| NL2Repo | 48.9 | 42.7 | 47.2 | 42.1 | 35.5 | 69.7 | 50.7 | 33.4 |
| DeepSWE | 46.2 | 18.0 | 18.0 | 20.0 | 8.0 | 58.0 | 70.0 | 10.0 |
| ProgramBench | 63.7 | 50.9 | – | – | 47.8 | 71.9 | 70.8 | 39.5 |
| Terminal-Bench 2.1 (Terminus-2) | 81.0 | 63.5 | 75.0 | 65.0 | 64.0 | 85.0 | 84.0 | 74.0 |
| Terminal-Bench 2.1 (best harness) | 82.7 (Claude Code) | 69.0 (Claude Code) | – | – | – | 78.9 (Claude Code) | 83.4 (Codex) | 70.7 (Gemini CLI) |
| FrontierSWE Dominance | 74.4 | 30.5 | – | – | 29.0 | 75.1 | 72.6 | 39.6 |
| PostTrainBench | 34.3 | 20.1 | – | – | – | 37.2 | 28.4 | 21.6 |
| SWE-Marathon | 13.0 | 1.0 | – | – | – | 26.0 | 12.0 | 4.0 |
| Agentic | ||||||||
| MCP-Atlas (public set) | 76.8 | 71.8 | 76.4 | 74.2 | 73.6 | 77.8 | 75.3 | 69.2 |
| Tool-Decathlon | 48.2 | 40.7 | – | – | 52.8 | 59.9 | 55.6 | 48.8 |
* score on the full set.
Published GLM model benchmark results, 17 evaluations across 8 models.
The evaluation settings below are the ones published alongside the results. They are worth reading before quoting any figure, because context length, sampling parameters, harness version and timeout budget all move scores materially — and because several of these evaluations were conducted by third parties rather than by the model authors.
Sampling at temperature 1.0 and top-p 0.95, with a maximum generation length of 163,840 tokens. The text-only subset is reported by default; results marked with an asterisk come from the full set. AIME, HMMT and IMOAnswerBench use a fixed answer-format system prompt, with GPT-5.5 (medium) as the judge model. HLE with tools uses a 300,000-token context and no context-management strategy.
Run with OpenHands using a tailored instruction prompt, at temperature 1, top-p 1 and a 32K new-token cap, inside a 400K context window.
Evaluated at temperature 1.0, top-p 1.0 and a 48K new-token cap under a 400K context. Rule-based and LLM-based judging is applied to block malicious behaviour such as unauthorised package installation or network calls.
Run with the official evaluation framework and the mini-swe-agent harness at temperature 1.0, top-p 1.0 and a two-hour timeout under a 400K context. Each task runs in an isolated container with 2 CPUs, 8 GB of RAM and no internet access.
200 instances evaluated with Claude Code 2.1.156 at temperature 1.0, top-p 1.0, a 64,000-token cap, up to 2,000 turns, a six-hour per-sample timeout and maximum reasoning effort, under a 400K context. Each instance runs in a 4-CPU, 8 GB sandbox with internet access disabled.
Evaluated with the Terminus-2 framework using JSON parsing, a four-hour timeout, temperature 1.0, top-p 1.0, a 48K new-token cap and up to 500 episodes, under a 256K context window. Resources are capped at 4 CPUs and 8 GB of RAM.
Evaluated in Claude Code 2.1.167 at temperature 1.0, top-p 0.95 and 131,072 new tokens, with the CLI's 64K output cap bypassed through a transparent proxy. Wall-clock limits are removed while per-task CPU and memory constraints are preserved. Scores are averaged over five runs.
All models evaluated in thinking mode on the 500-task public subset with a ten-minute timeout per task, using Gemini-3.0-Pro as the judge model.
Run through the official evaluation service with the maximum token budget set to 128K.
Conducted by Proximal, not by the model authors, using a 1M context length, the maximum effort level and 128K maximum output tokens. The dominance score is reported as of 16 June 2026.
Conducted by PostTrainBench, not by the model authors, using a 1M context length, the maximum effort level and 128K maximum output tokens.
Conducted by Abundant AI, not by the model authors, using a 1M context length, the maximum effort level and 128K maximum output tokens.
Read the GLM model benchmarks, then put a real model to work. The playground above is free and needs nothing from you; if you want a fuller AI toolkit, our partner's free tier starts here.