Model comparison

GLM-5.3-Flash vs Qwen3.6 Max Preview

GLM-5.3-Flash and Qwen3.6 Max Preview score almost the same on the Noometry Index (51.8 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

GLM-5.3-Flash Z.ai (Zhipu)

51.8

Rank #41 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 20 benchmarks with published results for both. GLM-5.3-Flash scores higher in 7 categories and Qwen3.6 Max Preview in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GLM-5.3-Flash leads 48.0 to 41.7.
  • The biggest single-benchmark swing is Mystery Game Puzzles: 8% for GLM-5.3-Flash and 19% for Qwen3.6 Max Preview.
  • GLM-5.3-Flash is cheaper at $0.15 / $0.50 per million input/output tokens, against $1.30 / $7.80 for Qwen3.6 Max Preview.
  • GLM-5.3-Flash accepts more context: 1M tokens versus 262K.
  • GLM-5.3-Flash has downloadable open weights; the other is API-only.

Side by side

GLM-5.3-Flash and Qwen3.6 Max Preview specifications
GLM-5.3-FlashQwen3.6 Max Preview
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index51.851.5
Released2026-08-202026-04-20
WeightsOpenProprietary
Context window1M262K
Max output131K66K
Input $ / M tokens$0.15$1.30
Output $ / M tokens$0.50$7.80
Results tracked4029

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3-Flash leads

GLM-5.3-Flash: 53.1 (#31), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
LMArena WebDev16091482
LMArena Coding15081471
SWE-bench Verified—76.7%
DeepSWE63.4%—
FrontierCode31.8%—
CursorBench36.8%—
FrontierSWE18.1%—
SciCode51.6%—
ALE-Bench303.55—

Agentic & Tool Use Not comparable

GLM-5.3-Flash: 34.2 (#47), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
APEX-Agents52.8%—
GDP.pdf14%—
Vending-Bench 2—4,254

Reasoning GLM-5.3-Flash leads

GLM-5.3-Flash: 48.0 (#42), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
Chess Puzzles14%20%
LMArena Hard Prompts14911457
Mystery Game Puzzles8%19%
Epoch Capabilities Index151.88149.24
ARC-AGI-265.8%—
SimpleBench—63%
NYT Connections (extended)—74.1%
ARC-AGI-191%—
CritPt15.4%—
DTBench—87.2%
LMCA—42.5%
Surface Evolver Bench52.5%—
Bench to the Future 30.15—

Math Too close to call

GLM-5.3-Flash: 53.3 (#47), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
OTIS Mock AIME 2024-202593.9%91.1%
LMArena Math15001465
FrontierMath (Tiers 1-3)55.8%—
FrontierMath Tier 417.1%—
ProofBench21%—
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Too close to call

GLM-5.3-Flash: 58.4 (#36), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
GPQA Diamond90.2%87.4%
LMArena Expert15131478
SimpleQA Verified—52%

Multimodal Not comparable

GLM-5.3-Flash: 42.8 (#27), Qwen3.6 Max Preview: —

Multimodal benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
LMArena Vision1296—

Multilingual GLM-5.3-Flash leads

GLM-5.3-Flash: 56.0 (#25), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
LMArena Non-English14621437
LMArena Chinese15271487
LMArena French14961449
LMArena Russian14691445
LMArena Spanish14711454
LMArena German1470—
LMArena Japanese1429—
LMArena Korean1446—

Instruction Following GLM-5.3-Flash leads

GLM-5.3-Flash: 77.5 (#20), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
LMArena Instruction Following14781438

Long Context Too close to call

GLM-5.3-Flash: 45.4 (#39), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
LMArena Longer Query14821457

Writing & Preference GLM-5.3-Flash leads

GLM-5.3-Flash: 65.3 (#50), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkGLM-5.3-FlashQwen3.6 Max Preview
LMArena Text14711447
LMArena Creative Writing14421435
LMArena Multi-Turn14671456

Frequently asked questions

Is GLM-5.3-Flash better than Qwen3.6 Max Preview?

GLM-5.3-Flash and Qwen3.6 Max Preview score almost the same on the Noometry Index (51.8 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.3-Flash or Qwen3.6 Max Preview?

GLM-5.3-Flash is cheaper. It lists at $0.15 per million input tokens and $0.50 per million output tokens; Qwen3.6 Max Preview lists at $1.30 and $7.80.

Is GLM-5.3-Flash or Qwen3.6 Max Preview better for coding?

GLM-5.3-Flash scores higher on coding benchmarks: 53.1 versus 48.7 in the Noometry coding category.

Which has the bigger context window?

GLM-5.3-Flash does, with 1M tokens against 262K.

How many benchmarks do GLM-5.3-Flash and Qwen3.6 Max Preview share?

20 benchmarks have published results for both models. GLM-5.3-Flash has 40 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper