Model comparison

GPT-5 vs Qwen3.6 Max Preview

GPT-5 and Qwen3.6 Max Preview score almost the same on the Noometry Index (50.9 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

GPT-5 OpenAI

50.9

Rank #45 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 27 benchmarks with published results for both. GPT-5 scores higher in 3 categories and Qwen3.6 Max Preview in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in long context, where GPT-5 leads 69.5 to 44.6.
  • The biggest single-benchmark swing is Chess Puzzles: 37% for GPT-5 and 20% for Qwen3.6 Max Preview.
  • Qwen3.6 Max Preview is cheaper at $1.30 / $7.80 per million input/output tokens, against $1.25 / $10 for GPT-5.
  • GPT-5 accepts more context: 400K tokens versus 262K.

Side by side

GPT-5 and Qwen3.6 Max Preview specifications
GPT-5Qwen3.6 Max Preview
ProviderOpenAIAlibaba (Qwen)
Noometry Index50.951.5
Released2025-08-072026-04-20
WeightsProprietaryProprietary
Context window400K262K
Max output128K66K
Input $ / M tokens$1.25$1.30
Output $ / M tokens$10$7.80
Results tracked6929

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5 leads

GPT-5: 50.3 (#47), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
SWE-bench Verified73.6%76.7%
LMArena WebDev14181482
LMArena Coding14361471
SWE-bench Verified (bash only)65%—
Aider Polyglot88%—
SciCode42.9%—
GSO6.9%—
WeirdML60.7%—
ALE-Bench1,162—
AlgoTune1.67—

Agentic & Tool Use Not comparable

GPT-5: 33.1 (#56), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
Terminal-Bench49.6%—
GDPval34.8%—
Remote Labor Index1.7%—
DeepResearch Bench49.6%—
BALROG32.8%—
LMArena Search1133—
METR Time Horizons69.6%—
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

GPT-5: 38.3 (#64), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
SimpleBench56.7%63%
Chess Puzzles37%20%
LMArena Hard Prompts14161457
Mystery Game Puzzles23%19%
DTBench90.7%87.2%
LMCA40%42.5%
Epoch Capabilities Index150149.24
ARC-AGI-29.9%—
Kagi LLM Benchmark72.7%—
NYT Connections (extended)—74.1%
ARC-AGI-165.7%—
CritPt12.6%—
EnigmaEval10.5%—
EBR-Bench12.7%—
ForecastBench61.4—

Math Too close to call

GPT-5: 55.0 (#44), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
OTIS Mock AIME 2024-202591.4%91.1%
LMArena Math14071465
FrontierMath (Feb 2025 set)32.4%23.1%
FrontierMath Tier 4 (v1)12.5%4.2%
FrontierMath (Tiers 1-3)55.4%—
FrontierMath Tier 422%—
ProofBench18%—
Omni-MATH64.7%—
MATH Level 598.1%—

Knowledge Qwen3.6 Max Preview leads

GPT-5: 56.6 (#43), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
GPQA Diamond86.2%87.4%
SimpleQA Verified50.1%52%
LMArena Expert14191478
Humanity's Last Exam25.3%—
MMLU-Pro86.3%—
Confabulations10.3%—
Vectara Hallucination Rate14.7%—
GPQA (HELM)79.2%—

Multimodal Not comparable

GPT-5: 46.8 (#13), Qwen3.6 Max Preview: —

Multimodal benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
LMArena Vision1232—
GeoBench81%—
VPCT66%—

Multilingual Qwen3.6 Max Preview leads

GPT-5: 51.4 (#110), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
LMArena Non-English13971437
LMArena Chinese14221487
LMArena French14101449
LMArena Russian14061445
LMArena Spanish13991454
LMArena German1416—
LMArena Japanese1409—
LMArena Korean1360—

Instruction Following Qwen3.6 Max Preview leads

GPT-5: 73.8 (#113), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
LMArena Instruction Following13881438
IFEval87.5%—

Long Context GPT-5 leads

GPT-5: 69.5 (#2), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
LMArena Longer Query13991457
Fiction.LiveBench97.2%—

Writing & Preference Too close to call

GPT-5: 63.4 (#65), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkGPT-5Qwen3.6 Max Preview
LMArena Text14061447
LMArena Creative Writing13651435
LMArena Multi-Turn14261456
Short-Story Creative Writing86%—
EQ-Bench Creative Writing1627—
WildBench85.7%—

Frequently asked questions

Is GPT-5 better than Qwen3.6 Max Preview?

GPT-5 and Qwen3.6 Max Preview score almost the same on the Noometry Index (50.9 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 or Qwen3.6 Max Preview?

Qwen3.6 Max Preview is cheaper. It lists at $1.30 per million input tokens and $7.80 per million output tokens; GPT-5 lists at $1.25 and $10.

Is GPT-5 or Qwen3.6 Max Preview better for coding?

GPT-5 scores higher on coding benchmarks: 50.3 versus 48.7 in the Noometry coding category.

Which has the bigger context window?

GPT-5 does, with 400K tokens against 262K.

How many benchmarks do GPT-5 and Qwen3.6 Max Preview share?

27 benchmarks have published results for both models. GPT-5 has 69 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper