Model comparison

o1 vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 40.9 on the Noometry Index.

Last verified . 27 shared benchmarks.

o1 OpenAI

40.9

Rank #143 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 27 benchmarks with published results for both. o1 scores higher in 3 categories and Qwen3 Max in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o1 leads 50.3 to 41.6.
  • The biggest single-benchmark swing is Fiction.LiveBench: 83.3% for o1 and 66.7% for Qwen3 Max.
  • Qwen3 Max is cheaper at $1.20 / $6 per million input/output tokens, against $15 / $60 for o1.
  • Qwen3 Max accepts more context: 262K tokens versus 200K.

Side by side

o1 and Qwen3 Max specifications
o1Qwen3 Max
ProviderOpenAIAlibaba (Qwen)
Noometry Index40.943.7
Released2024-09-122025-09-23
WeightsProprietaryProprietary
Context window200K262K
Max output100K66K
Input $ / M tokens$15$1.20
Output $ / M tokens$60$6
Results tracked5233

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

o1: 46.1 (#70), Qwen3 Max: 43.0 (#93)

Coding benchmarks
Benchmarko1Qwen3 Max
LMArena Coding13671456
Aider Polyglot61.7%—
WeirdML47.6%—
LiveBench Coding69.7%—
CadEval56%—
ALE-Bench—370.45
HumanEval+89%—
MBPP+80.2%—

Agentic & Tool Use Not comparable

o1: 24.6 (#117), Qwen3 Max: —

Agentic & Tool Use benchmarks
Benchmarko1Qwen3 Max
Cybench10%—
METR Time Horizons51.1%—
Vending-Bench 2—71.56

Reasoning o1 leads

o1: 27.9 (#111), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
Benchmarko1Qwen3 Max
Chess Puzzles15%4%
LMArena Hard Prompts13711448
DTBench74.7%82.1%
LMCA22.3%28.3%
Epoch Capabilities Index141.91142.38
SimpleBench41.7%—
Kagi LLM Benchmark—72.5%
NYT Connections (extended)—30.1%
ARC-AGI-130.7%—
EnigmaEval5.7%—
LiveBench Reasoning91.6%—
Mystery Game Puzzles—5%
LiveBench Data Analysis65.5%—
LiveBench75.7%—

Math Qwen3 Max leads

o1: 36.1 (#175), Qwen3 Max: 38.7 (#131)

Math benchmarks
Benchmarko1Qwen3 Max
FrontierMath (Tiers 1-3)14.7%18.9%
OTIS Mock AIME 2024-202573.3%73.3%
LMArena Math13881446
MATH Level 594.7%97.1%
LiveBench Math80.3%—
FrontierMath (Feb 2025 set)9.3%—

Knowledge Qwen3 Max leads

o1: 41.5 (#110), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
Benchmarko1Qwen3 Max
GPQA Diamond76.8%72.6%
SimpleQA Verified41.1%48.7%
LMArena Expert13611455
Humanity's Last Exam8%—
Confabulations11.7%—

Multimodal Not comparable

o1: 34.2 (#93), Qwen3 Max: —

Multimodal benchmarks
Benchmarko1Qwen3 Max
LMArena Vision1168—
GeoBench80%—
VPCT37%—
SpatialViz-Bench41.4%—

Multilingual Qwen3 Max leads

o1: 48.6 (#142), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
Benchmarko1Qwen3 Max
LMArena Non-English13581429
LMArena Chinese13941478
LMArena French13441449
LMArena German13371463
LMArena Japanese13461397
LMArena Korean13961399
LMArena Russian13561428
LMArena Spanish13451462

Instruction Following Too close to call

o1: 74.8 (#86), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
Benchmarko1Qwen3 Max
LMArena Instruction Following13671419
LiveBench Instruction Following81.5%—

Long Context o1 leads

o1: 50.3 (#9), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
Benchmarko1Qwen3 Max
Fiction.LiveBench83.3%66.7%
LMArena Longer Query13781438
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

o1: 55.6 (#144), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
Benchmarko1Qwen3 Max
LMArena Text13661439
LMArena Creative Writing13481402
LMArena Multi-Turn13691446
Short-Story Creative Writing70.2%—
LiveBench Language65.4%—

Frequently asked questions

Is o1 better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 40.9 on the Noometry Index.

Which is cheaper, o1 or Qwen3 Max?

Qwen3 Max is cheaper. It lists at $1.20 per million input tokens and $6 per million output tokens; o1 lists at $15 and $60.

Is o1 or Qwen3 Max better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 43.0 in the Noometry coding category.

Which has the bigger context window?

Qwen3 Max does, with 262K tokens against 200K.

How many benchmarks do o1 and Qwen3 Max share?

27 benchmarks have published results for both models. o1 has 52 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper