Model comparison

Qwen3.8 Max vs Step 3

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 40.5 on the Noometry Index.

Last verified . 16 shared benchmarks.

Qwen3.8 Max Alibaba (Qwen)

56.8

Rank #22 Confirmed

Step 3 StepFun

40.5

Rank #149 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Qwen3.8 Max scores higher in 9 categories and Step 3 in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.8 Max leads 73.2 to 37.6.
  • Step 3 has downloadable open weights; the other is API-only.

Side by side

Qwen3.8 Max and Step 3 specifications
Qwen3.8 MaxStep 3
ProviderAlibaba (Qwen)StepFun
Noometry Index56.840.5
Released2026-08-02—
WeightsProprietaryOpen
Context window1M—
Max output131K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3917

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.8 Max leads

Qwen3.8 Max: 53.5 (#29), Step 3: 40.1 (#147)

Coding benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Coding15021367
DeepSWE57.5%—
LMArena WebDev1674—
FrontierSWE17.8%—
SciCode53.2%—

Agentic & Tool Use Not comparable

Qwen3.8 Max: 45.4 (#14), Step 3: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.8 MaxStep 3
APEX-Agents63.3%—
τ²-bench Banking55.1%—
GDP.pdf23.2%—

Reasoning Qwen3.8 Max leads

Qwen3.8 Max: 54.4 (#26), Step 3: 28.4 (#105)

Reasoning benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Hard Prompts14961355
Kagi LLM Benchmark—62.3%
NYT Connections (extended)88.3%—
CritPt20%—
Chess Puzzles40%—
Mystery Game Puzzles38%—
DTBench92%—
LMCA46.2%—
Epoch Capabilities Index156.41—

Math Qwen3.8 Max leads

Qwen3.8 Max: 73.2 (#20), Step 3: 37.6 (#148)

Math benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Math14991366
FrontierMath (Tiers 1-3)74.7%—
FrontierMath Tier 446.3%—
OTIS Mock AIME 2024-2025100%—
ProofBench58%—

Knowledge Qwen3.8 Max leads

Qwen3.8 Max: 61.7 (#27), Step 3: 36.8 (#164)

Knowledge benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Expert15071333
GPQA Diamond92.7%—
SimpleQA Verified47.3%—

Multimodal Qwen3.8 Max leads

Qwen3.8 Max: 37.2 (#75), Step 3: 35.5 (#86)

Multimodal benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Vision13141177
Furniture Assembly20%—

Multilingual Qwen3.8 Max leads

Qwen3.8 Max: 56.7 (#18), Step 3: 46.3 (#159)

Multilingual benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Non-English14721327
LMArena Chinese15381397
LMArena German14831371
LMArena Korean14611269
LMArena Russian14811331
LMArena Spanish14921371
LMArena French1503—
LMArena Japanese1467—

Instruction Following Qwen3.8 Max leads

Qwen3.8 Max: 77.6 (#17), Step 3: 70.4 (#164)

Instruction Following benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Instruction Following14791332

Long Context Qwen3.8 Max leads

Qwen3.8 Max: 45.6 (#31), Step 3: 40.3 (#157)

Long Context benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Longer Query14891326

Writing & Preference Qwen3.8 Max leads

Qwen3.8 Max: 67.1 (#30), Step 3: 54.3 (#151)

Writing & Preference benchmarks
BenchmarkQwen3.8 MaxStep 3
LMArena Text14831350
LMArena Creative Writing14791321
LMArena Multi-Turn14891341

Frequently asked questions

Is Qwen3.8 Max better than Step 3?

Qwen3.8 Max is the stronger model overall, scoring 56.8 to 40.5 on the Noometry Index.

Is Qwen3.8 Max or Step 3 better for coding?

Qwen3.8 Max scores higher on coding benchmarks: 53.5 versus 40.1 in the Noometry coding category.

How many benchmarks do Qwen3.8 Max and Step 3 share?

16 benchmarks have published results for both models. Qwen3.8 Max has 39 scored results on Noometry and Step 3 has 17.

Related comparisons

Go deeper