Model comparison

GLM-5 vs Qwen3.7 Plus

GLM-5 and Qwen3.7 Plus score almost the same on the Noometry Index (46.1 vs 45.3), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

GLM-5 Z.ai (Zhipu)

46.1

Rank #66 Confirmed

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 22 benchmarks with published results for both. GLM-5 scores higher in 4 categories and Qwen3.7 Plus in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GLM-5 leads 49.0 to 36.6.
  • The biggest single-benchmark swing is Chess Puzzles: 10% for GLM-5 and 24% for Qwen3.7 Plus.
  • Qwen3.7 Plus is cheaper at $0.40 / $1.60 per million input/output tokens, against $1 / $3.20 for GLM-5.
  • Qwen3.7 Plus accepts more context: 1M tokens versus 205K.
  • GLM-5 has downloadable open weights; the other is API-only.

Side by side

GLM-5 and Qwen3.7 Plus specifications
GLM-5Qwen3.7 Plus
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index46.145.3
Released2026-02-112026-06-02
WeightsOpenProprietary
Context window205K1M
Max output131K131K
Input $ / M tokens$1$0.40
Output $ / M tokens$3.20$1.60
Results tracked4532

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5 leads

GLM-5: 49.0 (#52), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkGLM-5Qwen3.7 Plus
LMArena Coding14611473
SWE-bench Verified72.1%—
FrontierCode—10.2%
SWE-bench Verified (bash only)72.8%—
LMArena WebDev1434—
SWE-bench Multilingual69.7%—
SciCode—45.5%
WeirdML48.2%—
ALE-Bench765.62—

Agentic & Tool Use GLM-5 leads

GLM-5: 31.1 (#71), Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkGLM-5Qwen3.7 Plus
Terminal-Bench52.4%—
OSWorld 2.0—2.8%
τ²-bench Airline82.5%—
τ²-bench Banking9.8%—
τ²-bench Retail73.7%—
τ²-bench Telecom86.8%—
Vending-Bench 24,432—

Reasoning Qwen3.7 Plus leads

GLM-5: 27.6 (#116), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkGLM-5Qwen3.7 Plus
NYT Connections (extended)74.8%74.8%
Chess Puzzles10%24%
LMArena Hard Prompts14521460
Epoch Capabilities Index145.83147.37
ARC-AGI-24.9%—
SimpleBench53.2%—
Kagi LLM Benchmark75%—
ARC-AGI-144.7%—
CritPt—9.1%
Mystery Game Puzzles—17%
DTBench—84%
LMCA—37.6%
ForecastBench61—

Math Qwen3.7 Plus leads

GLM-5: 46.4 (#71), Qwen3.7 Plus: 50.5 (#56)

Knowledge Qwen3.7 Plus leads

GLM-5: 52.3 (#64), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkGLM-5Qwen3.7 Plus
GPQA Diamond87.8%87.9%
LMArena Expert14541467
Vectara Hallucination Rate10.1%—

Multimodal Not comparable

GLM-5: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkGLM-5Qwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Qwen3.7 Plus leads

GLM-5: 53.7 (#58), Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkGLM-5Qwen3.7 Plus
LMArena Non-English14301445
LMArena Chinese15111510
LMArena French14551473
LMArena German14451471
LMArena Japanese14161413
LMArena Korean14231415
LMArena Russian14361457
LMArena Spanish14541457

Instruction Following Too close to call

GLM-5: 75.2 (#67), Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkGLM-5Qwen3.7 Plus
LMArena Instruction Following14281440

Long Context Too close to call

GLM-5: 44.7 (#60), Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkGLM-5Qwen3.7 Plus
LMArena Longer Query14461455
CL-bench18.7%—

Writing & Preference GLM-5 leads

GLM-5: 66.0 (#38), Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkGLM-5Qwen3.7 Plus
LMArena Text14461455
LMArena Creative Writing14391439
LMArena Multi-Turn14561460
EQ-Bench Creative Writing1601—

Frequently asked questions

Is GLM-5 better than Qwen3.7 Plus?

GLM-5 and Qwen3.7 Plus score almost the same on the Noometry Index (46.1 vs 45.3), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5 or Qwen3.7 Plus?

Qwen3.7 Plus is cheaper. It lists at $0.40 per million input tokens and $1.60 per million output tokens; GLM-5 lists at $1 and $3.20.

Is GLM-5 or Qwen3.7 Plus better for coding?

GLM-5 scores higher on coding benchmarks: 49.0 versus 36.6 in the Noometry coding category.

Which has the bigger context window?

Qwen3.7 Plus does, with 1M tokens against 205K.

How many benchmarks do GLM-5 and Qwen3.7 Plus share?

22 benchmarks have published results for both models. GLM-5 has 45 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper