Model comparison

GLM-5.3-Flash vs Qwen3.7 Max

GLM-5.3-Flash and Qwen3.7 Max score almost the same on the Noometry Index (51.8 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

GLM-5.3-Flash Z.ai (Zhipu)

51.8

Rank #41 Confirmed

Qwen3.7 Max Alibaba (Qwen)

51.5

Rank #42 Confirmed

Summary

  • They share 24 benchmarks with published results for both. GLM-5.3-Flash scores higher in 4 categories and Qwen3.7 Max in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GLM-5.3-Flash leads 34.2 to 22.1.
  • The biggest single-benchmark swing is Mystery Game Puzzles: 8% for GLM-5.3-Flash and 32% for Qwen3.7 Max.
  • GLM-5.3-Flash is cheaper at $0.15 / $0.50 per million input/output tokens, against $2.50 / $7.50 for Qwen3.7 Max.
  • GLM-5.3-Flash has downloadable open weights; the other is API-only.

Side by side

GLM-5.3-Flash and Qwen3.7 Max specifications
GLM-5.3-FlashQwen3.7 Max
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index51.851.5
Released2026-08-202026-05-19
WeightsOpenProprietary
Context window1M1M
Max output131K131K
Input $ / M tokens$0.15$2.50
Output $ / M tokens$0.50$7.50
Results tracked4033

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3-Flash leads

GLM-5.3-Flash: 53.1 (#31), Qwen3.7 Max: 50.4 (#45)

Coding benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
LMArena WebDev16091515
SciCode51.6%48.8%
LMArena Coding15081498
ALE-Bench303.551,189
SWE-bench Verified—77.3%
DeepSWE63.4%—
FrontierCode31.8%—
CursorBench36.8%—
FrontierSWE18.1%—

Agentic & Tool Use GLM-5.3-Flash leads

GLM-5.3-Flash: 34.2 (#47), Qwen3.7 Max: 22.1 (#135)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
APEX-Agents52.8%—
GBAEval—0.4%
GDP.pdf14%—

Reasoning Qwen3.7 Max leads

GLM-5.3-Flash: 48.0 (#42), Qwen3.7 Max: 49.2 (#38)

Reasoning benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
CritPt15.4%13.4%
Chess Puzzles14%19%
LMArena Hard Prompts14911483
Mystery Game Puzzles8%32%
Epoch Capabilities Index151.88153.68
ARC-AGI-265.8%—
SimpleBench—70.4%
NYT Connections (extended)—85.1%
ARC-AGI-191%—
EBR-Bench—9.5%
DTBench—92.3%
LMCA—44%
Surface Evolver Bench52.5%—
Bench to the Future 30.15—

Math Qwen3.7 Max leads

GLM-5.3-Flash: 53.3 (#47), Qwen3.7 Max: 62.4 (#32)

Math benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
FrontierMath (Tiers 1-3)55.8%64.6%
FrontierMath Tier 417.1%34.1%
OTIS Mock AIME 2024-202593.9%95.6%
ProofBench21%26%
LMArena Math15001490

Knowledge Qwen3.7 Max leads

GLM-5.3-Flash: 58.4 (#36), Qwen3.7 Max: 61.6 (#28)

Knowledge benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
GPQA Diamond90.2%90.9%
LMArena Expert15131488
SimpleQA Verified—55.8%

Multimodal Not comparable

GLM-5.3-Flash: 42.8 (#27), Qwen3.7 Max: —

Multimodal benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
LMArena Vision1296—

Multilingual Too close to call

GLM-5.3-Flash: 56.0 (#25), Qwen3.7 Max: 56.9 (#15)

Multilingual benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
LMArena Non-English14621474
LMArena Chinese15271530
LMArena Russian14691484
LMArena French1496—
LMArena German1470—
LMArena Japanese1429—
LMArena Korean1446—
LMArena Spanish1471—

Instruction Following Too close to call

GLM-5.3-Flash: 77.5 (#20), Qwen3.7 Max: 76.7 (#38)

Instruction Following benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
LMArena Instruction Following14781460

Long Context Too close to call

GLM-5.3-Flash: 45.4 (#39), Qwen3.7 Max: 45.4 (#40)

Long Context benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
LMArena Longer Query14821482

Writing & Preference Too close to call

GLM-5.3-Flash: 65.3 (#50), Qwen3.7 Max: 65.0 (#54)

Writing & Preference benchmarks
BenchmarkGLM-5.3-FlashQwen3.7 Max
LMArena Text14711476
LMArena Creative Writing14421449
LMArena Multi-Turn14671481
EQ-Bench 4—1110

Frequently asked questions

Is GLM-5.3-Flash better than Qwen3.7 Max?

GLM-5.3-Flash and Qwen3.7 Max score almost the same on the Noometry Index (51.8 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.3-Flash or Qwen3.7 Max?

GLM-5.3-Flash is cheaper. It lists at $0.15 per million input tokens and $0.50 per million output tokens; Qwen3.7 Max lists at $2.50 and $7.50.

Is GLM-5.3-Flash or Qwen3.7 Max better for coding?

GLM-5.3-Flash scores higher on coding benchmarks: 53.1 versus 50.4 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do GLM-5.3-Flash and Qwen3.7 Max share?

24 benchmarks have published results for both models. GLM-5.3-Flash has 40 scored results on Noometry and Qwen3.7 Max has 33.

Related comparisons

Go deeper