Model comparison

GPT-5.2 Pro vs Qwen3.7 Max

GPT-5.2 Pro and Qwen3.7 Max score almost the same on the Noometry Index (52.3 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 5 shared benchmarks.

GPT-5.2 Pro OpenAI

52.3

Rank #39 Reported

Qwen3.7 Max Alibaba (Qwen)

51.5

Rank #42 Confirmed

Summary

  • They share 5 benchmarks with published results for both. GPT-5.2 Pro scores higher in 2 categories and Qwen3.7 Max in 0 categories; 2 gaps are clear of the uncertainty.
  • The biggest single-benchmark swing is SimpleBench: 57.4% for GPT-5.2 Pro and 70.4% for Qwen3.7 Max.
  • Qwen3.7 Max is cheaper at $2.50 / $7.50 per million input/output tokens, against $21 / $168 for GPT-5.2 Pro.
  • Qwen3.7 Max accepts more context: 1M tokens versus 400K.

Side by side

GPT-5.2 Pro and Qwen3.7 Max specifications
GPT-5.2 ProQwen3.7 Max
ProviderOpenAIAlibaba (Qwen)
Noometry Index52.351.5
Released2025-12-112026-05-19
WeightsProprietaryProprietary
Context window400K1M
Max output128K131K
Input $ / M tokens$21$2.50
Output $ / M tokens$168$7.50
Results tracked833

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

GPT-5.2 Pro: —, Qwen3.7 Max: 50.4 (#45)

Coding benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
SWE-bench Verified—77.3%
LMArena WebDev—1515
SciCode—48.8%
LMArena Coding—1498
ALE-Bench—1,189

Agentic & Tool Use Not comparable

GPT-5.2 Pro: —, Qwen3.7 Max: 22.1 (#135)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
GBAEval—0.4%

Reasoning GPT-5.2 Pro leads

GPT-5.2 Pro: 51.5 (#33), Qwen3.7 Max: 49.2 (#38)

Reasoning benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
SimpleBench57.4%70.4%
NYT Connections (extended)79.3%85.1%
Epoch Capabilities Index155.4153.68
ARC-AGI-254.2%—
ARC-AGI-190.5%—
CritPt—13.4%
Chess Puzzles—19%
EBR-Bench—9.5%
LMArena Hard Prompts—1483
Mystery Game Puzzles—32%
DTBench—92.3%
LMCA—44%

Math GPT-5.2 Pro leads

GPT-5.2 Pro: 65.3 (#29), Qwen3.7 Max: 62.4 (#32)

Math benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
FrontierMath (Tiers 1-3)74%64.6%
FrontierMath Tier 446%34.1%
OTIS Mock AIME 2024-2025—95.6%
ProofBench—26%
LMArena Math—1490
FrontierMath Tier 4 (v1)31.3%—

Knowledge Not comparable

GPT-5.2 Pro: —, Qwen3.7 Max: 61.6 (#28)

Knowledge benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
GPQA Diamond—90.9%
SimpleQA Verified—55.8%
LMArena Expert—1488

Multilingual Not comparable

GPT-5.2 Pro: —, Qwen3.7 Max: 56.9 (#15)

Multilingual benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
LMArena Non-English—1474
LMArena Chinese—1530
LMArena Russian—1484

Instruction Following Not comparable

GPT-5.2 Pro: —, Qwen3.7 Max: 76.7 (#38)

Instruction Following benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
LMArena Instruction Following—1460

Long Context Not comparable

GPT-5.2 Pro: —, Qwen3.7 Max: 45.4 (#40)

Long Context benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
LMArena Longer Query—1482

Writing & Preference Not comparable

GPT-5.2 Pro: —, Qwen3.7 Max: 65.0 (#54)

Writing & Preference benchmarks
BenchmarkGPT-5.2 ProQwen3.7 Max
LMArena Text—1476
LMArena Creative Writing—1449
EQ-Bench 4—1110
LMArena Multi-Turn—1481

Frequently asked questions

Is GPT-5.2 Pro better than Qwen3.7 Max?

GPT-5.2 Pro and Qwen3.7 Max score almost the same on the Noometry Index (52.3 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.2 Pro or Qwen3.7 Max?

Qwen3.7 Max is cheaper. It lists at $2.50 per million input tokens and $7.50 per million output tokens; GPT-5.2 Pro lists at $21 and $168.

Which has the bigger context window?

Qwen3.7 Max does, with 1M tokens against 400K.

How many benchmarks do GPT-5.2 Pro and Qwen3.7 Max share?

5 benchmarks have published results for both models. GPT-5.2 Pro has 8 scored results on Noometry and Qwen3.7 Max has 33.

Related comparisons

Go deeper