Model comparison

GPT-5 vs Qwen3.7 Max

GPT-5 and Qwen3.7 Max score almost the same on the Noometry Index (50.9 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 30 shared benchmarks.

GPT-5 OpenAI

50.9

Rank #45 Confirmed

Qwen3.7 Max Alibaba (Qwen)

51.5

Rank #42 Confirmed

Summary

  • They share 30 benchmarks with published results for both. GPT-5 scores higher in 2 categories and Qwen3.7 Max in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where GPT-5 leads 69.5 to 45.4.
  • The biggest single-benchmark swing is Chess Puzzles: 37% for GPT-5 and 19% for Qwen3.7 Max.
  • GPT-5 is cheaper at $1.25 / $10 per million input/output tokens, against $2.50 / $7.50 for Qwen3.7 Max.
  • Qwen3.7 Max accepts more context: 1M tokens versus 400K.

Side by side

GPT-5 and Qwen3.7 Max specifications
GPT-5Qwen3.7 Max
ProviderOpenAIAlibaba (Qwen)
Noometry Index50.951.5
Released2025-08-072026-05-19
WeightsProprietaryProprietary
Context window400K1M
Max output128K131K
Input $ / M tokens$1.25$2.50
Output $ / M tokens$10$7.50
Results tracked6933

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-5: 50.3 (#47), Qwen3.7 Max: 50.4 (#45)

Coding benchmarks
BenchmarkGPT-5Qwen3.7 Max
SWE-bench Verified73.6%77.3%
LMArena WebDev14181515
SciCode42.9%48.8%
LMArena Coding14361498
ALE-Bench1,1621,189
SWE-bench Verified (bash only)65%—
Aider Polyglot88%—
GSO6.9%—
WeirdML60.7%—
AlgoTune1.67—

Agentic & Tool Use GPT-5 leads

GPT-5: 33.1 (#56), Qwen3.7 Max: 22.1 (#135)

Agentic & Tool Use benchmarks
BenchmarkGPT-5Qwen3.7 Max
Terminal-Bench49.6%—
GDPval34.8%—
Remote Labor Index1.7%—
DeepResearch Bench49.6%—
BALROG32.8%—
GBAEval—0.4%
LMArena Search1133—
METR Time Horizons69.6%—

Reasoning Qwen3.7 Max leads

GPT-5: 38.3 (#64), Qwen3.7 Max: 49.2 (#38)

Reasoning benchmarks
BenchmarkGPT-5Qwen3.7 Max
SimpleBench56.7%70.4%
CritPt12.6%13.4%
Chess Puzzles37%19%
EBR-Bench12.7%9.5%
LMArena Hard Prompts14161483
Mystery Game Puzzles23%32%
DTBench90.7%92.3%
LMCA40%44%
Epoch Capabilities Index150153.68
ARC-AGI-29.9%—
Kagi LLM Benchmark72.7%—
NYT Connections (extended)—85.1%
ARC-AGI-165.7%—
EnigmaEval10.5%—
ForecastBench61.4—

Math Qwen3.7 Max leads

GPT-5: 55.0 (#44), Qwen3.7 Max: 62.4 (#32)

Math benchmarks
BenchmarkGPT-5Qwen3.7 Max
FrontierMath (Tiers 1-3)55.4%64.6%
FrontierMath Tier 422%34.1%
OTIS Mock AIME 2024-202591.4%95.6%
ProofBench18%26%
LMArena Math14071490
Omni-MATH64.7%—
MATH Level 598.1%—
FrontierMath (Feb 2025 set)32.4%—
FrontierMath Tier 4 (v1)12.5%—

Knowledge Qwen3.7 Max leads

GPT-5: 56.6 (#43), Qwen3.7 Max: 61.6 (#28)

Knowledge benchmarks
BenchmarkGPT-5Qwen3.7 Max
GPQA Diamond86.2%90.9%
SimpleQA Verified50.1%55.8%
LMArena Expert14191488
Humanity's Last Exam25.3%—
MMLU-Pro86.3%—
Confabulations10.3%—
Vectara Hallucination Rate14.7%—
GPQA (HELM)79.2%—

Multimodal Not comparable

GPT-5: 46.8 (#13), Qwen3.7 Max: —

Multimodal benchmarks
BenchmarkGPT-5Qwen3.7 Max
LMArena Vision1232—
GeoBench81%—
VPCT66%—

Multilingual Qwen3.7 Max leads

GPT-5: 51.4 (#110), Qwen3.7 Max: 56.9 (#15)

Multilingual benchmarks
BenchmarkGPT-5Qwen3.7 Max
LMArena Non-English13971474
LMArena Chinese14221530
LMArena Russian14061484
LMArena French1410—
LMArena German1416—
LMArena Japanese1409—
LMArena Korean1360—
LMArena Spanish1399—

Instruction Following Qwen3.7 Max leads

GPT-5: 73.8 (#113), Qwen3.7 Max: 76.7 (#38)

Instruction Following benchmarks
BenchmarkGPT-5Qwen3.7 Max
LMArena Instruction Following13881460
IFEval87.5%—

Long Context GPT-5 leads

GPT-5: 69.5 (#2), Qwen3.7 Max: 45.4 (#40)

Long Context benchmarks
BenchmarkGPT-5Qwen3.7 Max
LMArena Longer Query13991482
Fiction.LiveBench97.2%—

Writing & Preference Qwen3.7 Max leads

GPT-5: 63.4 (#65), Qwen3.7 Max: 65.0 (#54)

Writing & Preference benchmarks
BenchmarkGPT-5Qwen3.7 Max
LMArena Text14061476
LMArena Creative Writing13651449
LMArena Multi-Turn14261481
Short-Story Creative Writing86%—
EQ-Bench Creative Writing1627—
WildBench85.7%—
EQ-Bench 4—1110

Frequently asked questions

Is GPT-5 better than Qwen3.7 Max?

GPT-5 and Qwen3.7 Max score almost the same on the Noometry Index (50.9 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 or Qwen3.7 Max?

GPT-5 is cheaper. It lists at $1.25 per million input tokens and $10 per million output tokens; Qwen3.7 Max lists at $2.50 and $7.50.

Is GPT-5 or Qwen3.7 Max better for coding?

They score almost the same on coding (50.3 vs 50.4); test both on your own repository before choosing.

Which has the bigger context window?

Qwen3.7 Max does, with 1M tokens against 400K.

How many benchmarks do GPT-5 and Qwen3.7 Max share?

30 benchmarks have published results for both models. GPT-5 has 69 scored results on Noometry and Qwen3.7 Max has 33.

Related comparisons

Go deeper