Model comparison

GLM-5.2 vs Qwen3.7 Max

GLM-5.2 and Qwen3.7 Max score almost the same on the Noometry Index (51.1 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 33 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Qwen3.7 Max Alibaba (Qwen)

51.5

Rank #42 Confirmed

Summary

  • They share 33 benchmarks with published results for both. GLM-5.2 scores higher in 4 categories and Qwen3.7 Max in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GLM-5.2 leads 32.4 to 22.1.
  • The biggest single-benchmark swing is SimpleQA Verified: 34.2% for GLM-5.2 and 55.8% for Qwen3.7 Max.
  • GLM-5.2 is cheaper at $1.40 / $4.40 per million input/output tokens, against $2.50 / $7.50 for Qwen3.7 Max.
  • GLM-5.2 has downloadable open weights; the other is API-only.

Side by side

GLM-5.2 and Qwen3.7 Max specifications
GLM-5.2Qwen3.7 Max
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index51.151.5
Released2026-06-132026-05-19
WeightsOpenProprietary
Context window1M1M
Max output131K131K
Input $ / M tokens$1.40$2.50
Output $ / M tokens$4.40$7.50
Results tracked5133

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-5.2: 51.3 (#41), Qwen3.7 Max: 50.4 (#45)

Coding benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
SWE-bench Verified78.7%77.3%
LMArena WebDev16031515
SciCode50.5%48.8%
LMArena Coding14851498
ALE-Bench1,0471,189
DeepSWE43.8%—
FrontierCode24.5%—
WeirdML70.1%—

Agentic & Tool Use GLM-5.2 leads

GLM-5.2: 32.4 (#63), Qwen3.7 Max: 22.1 (#135)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
GBAEval0%0.4%
APEX-Agents45.2%—
τ²-bench Banking37.1%—
PostTrainBench31.7%—
Vending-Bench 28,314—

Reasoning Qwen3.7 Max leads

GLM-5.2: 42.3 (#52), Qwen3.7 Max: 49.2 (#38)

Reasoning benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
SimpleBench58.8%70.4%
NYT Connections (extended)74.3%85.1%
CritPt20.9%13.4%
Chess Puzzles21%19%
EBR-Bench9.5%9.5%
LMArena Hard Prompts14801483
Mystery Game Puzzles19%32%
DTBench93.6%92.3%
LMCA45.8%44%
Epoch Capabilities Index151.78153.68
ARC-AGI-222.8%—
Kagi LLM Benchmark62.6%—
ARC-AGI-177%—
Surface Evolver Bench55.6%—

Math Qwen3.7 Max leads

GLM-5.2: 55.7 (#43), Qwen3.7 Max: 62.4 (#32)

Math benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
FrontierMath (Tiers 1-3)59.2%64.6%
FrontierMath Tier 429.3%34.1%
OTIS Mock AIME 2024-202586.4%95.6%
ProofBench35%26%
LMArena Math14821490
MathArena Final-Answer Competitions67.6%—

Knowledge Qwen3.7 Max leads

GLM-5.2: 57.1 (#40), Qwen3.7 Max: 61.6 (#28)

Knowledge benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
GPQA Diamond91.9%90.9%
SimpleQA Verified34.2%55.8%
LMArena Expert14861488

Multilingual Qwen3.7 Max leads

GLM-5.2: 55.8 (#26), Qwen3.7 Max: 56.9 (#15)

Multilingual benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
LMArena Non-English14591474
LMArena Chinese15191530
LMArena Russian14661484
LMArena French1479—
LMArena German1468—
LMArena Japanese1451—
LMArena Korean1445—
LMArena Spanish1477—

Instruction Following Too close to call

GLM-5.2: 76.9 (#34), Qwen3.7 Max: 76.7 (#38)

Instruction Following benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
LMArena Instruction Following14651460

Long Context Too close to call

GLM-5.2: 45.3 (#43), Qwen3.7 Max: 45.4 (#40)

Long Context benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
LMArena Longer Query14791482

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), Qwen3.7 Max: 65.0 (#54)

Writing & Preference benchmarks
BenchmarkGLM-5.2Qwen3.7 Max
LMArena Text14701476
LMArena Creative Writing14621449
EQ-Bench 412221110
LMArena Multi-Turn14691481
EQ-Bench Creative Writing1757—

Frequently asked questions

Is GLM-5.2 better than Qwen3.7 Max?

GLM-5.2 and Qwen3.7 Max score almost the same on the Noometry Index (51.1 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.2 or Qwen3.7 Max?

GLM-5.2 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; Qwen3.7 Max lists at $2.50 and $7.50.

Is GLM-5.2 or Qwen3.7 Max better for coding?

They score almost the same on coding (51.3 vs 50.4); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do GLM-5.2 and Qwen3.7 Max share?

33 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Qwen3.7 Max has 33.

Related comparisons

Go deeper