GLM vs Qwen3.7-Max pits two of the strongest open-weight releases against each other, and the result is closer than the DeepSeek comparison. Across the 13 published evaluations where both have a score, GLM-5.2 takes 10 and Qwen3.7-Max takes three — but Qwen's three are all in mathematics and knowledge-heavy reasoning, which makes the split easy to reason about.
Every figure here is a result as published by the model authors for each model, reproduced without adjustment. This site does not run the evaluations, and it does not extrapolate a score into a row where the authors published none — several benchmarks are simply absent from this comparison for that reason.
GLM-5.2 wins 10 of the 13 rows, and every coding row among them: SWE-bench Pro 62.1 to 60.6, NL2Repo 48.9 to 47.2, Terminal-Bench 2.1 81.0 to 75.0 and DeepSWE 46.2 to 18.0, which is the widest gap on the page. Qwen3.7-Max wins HLE at 41.4 to 40.5, HMMT Nov. 2025 at 95.0 to 94.4 and HMMT Feb. 2026 at 97.1 to 92.5. Two of those three margins are under a point. The honest reading is that these models are near-peers on knowledge and mathematics, and that GLM-5.2 pulls away as tasks become longer and more agentic.
Both models sit in the open-weight world, so the meaningful structural differences are context and control rather than licensing philosophy. GLM-5.2 publishes a 1,000,000-token window and three explicit effort levels; neither figure is part of the published comparison set for Qwen3.7-Max, and this table says so instead of filling the cell with a guess.
| Feature | GLM-5.2 | Qwen3.7-Max |
|---|---|---|
| Context window | 1,000,000 tokens | Not published |
| Openness | Open weights | Open-weight family |
| License | MIT license | See the vendor's release terms |
| Effort control | Explicit — Non-Thinking, High, Max | Not published |
| Self-hosting | Yes — transformers, vLLM, SGLang, xLLM, ktransformers | Vendor-dependent |
The coding block is a clean sweep for GLM-5.2, though three of the four margins are small — one to two points on SWE-bench Pro and NL2Repo. DeepSWE is the exception at 46.2 to 18.0, a gap large enough to suggest a real difference in long-horizon execution rather than measurement noise. In the reasoning block the models trade rows within a point of each other on most benchmarks, with GLM-5.2 clearly ahead on CritPt (20.9 to 13.4) and AIME 2026 (99.2 to 97.0). Counting wins overstates the distance here. Strip out every row decided by less than a point and the comparison reduces to three meaningful results: DeepSWE and CritPt to GLM-5.2, HMMT Feb. 2026 to Qwen3.7-Max. Two of those three reward sustained execution and one rewards mathematical depth, which is a fair summary of how these two open-weight models genuinely differ in day-to-day practice rather than on a scoreboard.
| Benchmark | GLM-5.2 | Qwen3.7-Max | Winner |
|---|---|---|---|
| Coding & long-horizon | |||
| SWE-bench Pro | 62.1 | 60.6 | GLM-5.2 |
| NL2Repo | 48.9 | 47.2 | GLM-5.2 |
| DeepSWE | 46.2 | 18.0 | GLM-5.2 |
| Terminal-Bench 2.1 (Terminus-2) | 81.0 | 75.0 | GLM-5.2 |
| Reasoning & agentic | |||
| MCP-Atlas (public set) | 76.8 | 76.4 | GLM-5.2 |
| HLE | 40.5 | 41.4 | Qwen3.7-Max |
| HLE w/ Tools | 54.7 | 53.5 | GLM-5.2 |
| CritPt | 20.9 | 13.4 | GLM-5.2 |
| AIME 2026 | 99.2 | 97.0 | GLM-5.2 |
| HMMT Nov. 2025 | 94.4 | 95.0 | Qwen3.7-Max |
| HMMT Feb. 2026 | 92.5 | 97.1 | Qwen3.7-Max |
| IMOAnswerBench | 91.0 | 90.0 | GLM-5.2 |
| GPQA-Diamond | 91.2 | 90.0 | GLM-5.2 |
* score on the full set.
All figures on this page are results as published by the model authors. glmmodel.com reports them; it does not run these evaluations.
See all 17 GLM benchmarks in detail →Choose the GLM model when the task has length to it. Its coding wins here are narrow on bounded problems and enormous on DeepSWE, which is the pattern you would expect from a model trained on long coding-agent trajectories with a 1,000,000-token window behind it.
Qwen3.7-Max is the stronger model on this table for mathematics and broad knowledge. Its 97.1 on HMMT Feb. 2026 is a genuine lead, and its HLE and HMMT Nov. 2025 results edge GLM-5.2 by fractions of a point. For a workload dominated by hard, self-contained questions, it is a serious open-weight alternative.
Read the GLM model benchmarks, then put a real model to work. The playground above is free and needs nothing from you; if you want a fuller AI toolkit, our partner's free tier starts here.