Model comparison

Qwen Max vs Qwen3.8 Max

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 34.7 on the Noometry Index.

Last verified . 19 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Qwen3.8 Max Alibaba (Qwen)

56.8

Rank #22 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Qwen Max scores higher in 0 categories and Qwen3.8 Max in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.8 Max leads 73.2 to 22.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 16.1% for Qwen Max and 100% for Qwen3.8 Max.
  • Qwen Max is cheaper at $1.60 / $6.40 per million input/output tokens, against $2 / $6 for Qwen3.8 Max.
  • Qwen3.8 Max accepts more context: 1M tokens versus 33K.

Side by side

Qwen Max and Qwen3.8 Max specifications
Qwen MaxQwen3.8 Max
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.756.8
Released2024-04-032026-08-02
WeightsProprietaryProprietary
Context window33K1M
Max output8K131K
Input $ / M tokens$1.60$2
Output $ / M tokens$6.40$6
Results tracked2339

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 Max leads

Qwen Max: 30.7 (#292), Qwen3.8 Max: 53.5 (#29)

Coding benchmarks
BenchmarkQwen MaxQwen3.8 Max
LMArena Coding12881502
DeepSWE—57.5%
Aider Polyglot21.8%—
LMArena WebDev—1674
FrontierSWE—17.8%
SciCode—53.2%

Agentic & Tool Use Not comparable

Qwen Max: —, Qwen3.8 Max: 45.4 (#14)

Agentic & Tool Use benchmarks
BenchmarkQwen MaxQwen3.8 Max
APEX-Agents—63.3%
τ²-bench Banking—55.1%
GDP.pdf—23.2%

Reasoning Qwen3.8 Max leads

Qwen Max: 25.1 (#151), Qwen3.8 Max: 54.4 (#26)

Reasoning benchmarks
BenchmarkQwen MaxQwen3.8 Max
LMArena Hard Prompts12691496
NYT Connections (extended)—88.3%
CritPt—20%
Chess Puzzles—40%
Mystery Game Puzzles—38%
DTBench—92%
LMCA—46.2%
Epoch Capabilities Index—156.41

Math Qwen3.8 Max leads

Qwen Max: 22.3 (#276), Qwen3.8 Max: 73.2 (#20)

Math benchmarks
BenchmarkQwen MaxQwen3.8 Max
OTIS Mock AIME 2024-202516.1%100%
LMArena Math12751499
FrontierMath (Tiers 1-3)—74.7%
FrontierMath Tier 4—46.3%
ProofBench—58%
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Qwen3.8 Max leads

Qwen Max: 30.3 (#228), Qwen3.8 Max: 61.7 (#27)

Knowledge benchmarks
BenchmarkQwen MaxQwen3.8 Max
GPQA Diamond56.1%92.7%
LMArena Expert12481507
SimpleQA Verified—47.3%

Multimodal Not comparable

Qwen Max: —, Qwen3.8 Max: 37.2 (#75)

Multimodal benchmarks
BenchmarkQwen MaxQwen3.8 Max
LMArena Vision—1314
Furniture Assembly—20%

Multilingual Qwen3.8 Max leads

Qwen Max: 41.8 (#202), Qwen3.8 Max: 56.7 (#18)

Multilingual benchmarks
BenchmarkQwen MaxQwen3.8 Max
LMArena Non-English12631472
LMArena Chinese12541538
LMArena French13301503
LMArena German12541483
LMArena Japanese12051467
LMArena Korean11421461
LMArena Russian12741481
LMArena Spanish12901492

Instruction Following Qwen3.8 Max leads

Qwen Max: 66.5 (#208), Qwen3.8 Max: 77.6 (#17)

Instruction Following benchmarks
BenchmarkQwen MaxQwen3.8 Max
LMArena Instruction Following12621479

Long Context Qwen3.8 Max leads

Qwen Max: 39.4 (#180), Qwen3.8 Max: 45.6 (#31)

Long Context benchmarks
BenchmarkQwen MaxQwen3.8 Max
LMArena Longer Query12881489
Fiction.LiveBench66.7%—

Writing & Preference Qwen3.8 Max leads

Qwen Max: 47.8 (#205), Qwen3.8 Max: 67.1 (#30)

Writing & Preference benchmarks
BenchmarkQwen MaxQwen3.8 Max
LMArena Text12821483
LMArena Creative Writing12481479
LMArena Multi-Turn12771489

Frequently asked questions

Is Qwen Max better than Qwen3.8 Max?

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 34.7 on the Noometry Index.

Which is cheaper, Qwen Max or Qwen3.8 Max?

Qwen Max is cheaper. It lists at $1.60 per million input tokens and $6.40 per million output tokens; Qwen3.8 Max lists at $2 and $6.

Is Qwen Max or Qwen3.8 Max better for coding?

Qwen3.8 Max scores higher on coding benchmarks: 53.5 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Qwen3.8 Max does, with 1M tokens against 33K.

How many benchmarks do Qwen Max and Qwen3.8 Max share?

19 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Qwen3.8 Max has 39.

Related comparisons

Go deeper