Model comparison

Claude 3.5 Sonnet vs Qwen Max

Claude 3.5 Sonnet and Qwen Max score almost the same on the Noometry Index (34.6 vs 34.7), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 5 categories and Qwen Max in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Claude 3.5 Sonnet leads 39.0 to 30.7.
  • The biggest single-benchmark swing is Aider Polyglot: 51.6% for Claude 3.5 Sonnet and 21.8% for Qwen Max.

Side by side

Claude 3.5 Sonnet and Qwen Max specifications
Claude 3.5 SonnetQwen Max
ProviderAnthropicAlibaba (Qwen)
Noometry Index34.634.7
Released2024-06-202024-04-03
WeightsProprietaryProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked6023

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
Aider Polyglot51.6%21.8%
LMArena Coding13421288
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Qwen Max leads

Claude 3.5 Sonnet: 23.1 (#183), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
LMArena Hard Prompts13051269
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Qwen Max leads

Claude 3.5 Sonnet: 19.2 (#288), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
OTIS Mock AIME 2024-20258.5%16.1%
LMArena Math13071275
MATH Level 556.9%67.2%
FrontierMath (Feb 2025 set)2.1%1%
Omni-MATH27.6%—
LiveBench Math52.3%—
FrontierMath Tier 4 (v1)0%—

Knowledge Qwen Max leads

Claude 3.5 Sonnet: 28.6 (#245), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
GPQA Diamond55.3%56.1%
LMArena Expert12651248
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Qwen Max: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 43.2 (#185), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
LMArena Non-English12831263
LMArena Chinese12721254
LMArena French13051330
LMArena German12971254
LMArena Japanese12341205
LMArena Korean12001142
LMArena Russian13061274
LMArena Spanish12901290

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
LMArena Instruction Following12971262
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
LMArena Longer Query13111288
Fiction.LiveBench—66.7%

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetQwen Max
LMArena Text12981282
LMArena Creative Writing12921248
LMArena Multi-Turn13261277
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Qwen Max?

Claude 3.5 Sonnet and Qwen Max score almost the same on the Noometry Index (34.6 vs 34.7), so choose on price, context window or the category you care about most.

Is Claude 3.5 Sonnet or Qwen Max better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 30.7 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Qwen Max share?

22 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper