Model comparison

Claude Fable 5.1 vs Qwen3 Max

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 43.7 on the Noometry Index. Qwen3 Max costs 8.3× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Last verified . 28 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Claude Fable 5.1 scores higher in 8 categories and Qwen3 Max in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Fable 5.1 leads 76.7 to 22.6.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 90.2% for Claude Fable 5.1 and 18.9% for Qwen3 Max.
  • Qwen3 Max is cheaper at $1.20 / $6 per million input/output tokens, against $10 / $50 for Claude Fable 5.1.
  • Claude Fable 5.1 accepts more context: 1M tokens versus 262K.

Side by side

Claude Fable 5.1 and Qwen3 Max specifications
Claude Fable 5.1Qwen3 Max
ProviderAnthropicAlibaba (Qwen)
Noometry Index69.043.7
Released2026-09-012025-09-23
WeightsProprietaryProprietary
Context window1M262K
Max output128K66K
Input $ / M tokens$10$1.20
Output $ / M tokens$50$6
Results tracked5233

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
LMArena Coding15281456
ALE-Bench2,143370.45
FrontierCode50.9%—
CursorBench51.8%—
LMArena WebDev1744—
FrontierSWE56.3%—
SciCode63.1%—
GSO88.2%—
WeirdML92.9%—
MirrorCode73.3%—

Agentic & Tool Use Not comparable

Claude Fable 5.1: 50.7 (#5), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
Vending-Bench 25,42271.56
APEX-Agents68.6%—
Remote Labor Index17.9%—
GDP.pdf29.6%—

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
NYT Connections (extended)90%30.1%
Chess Puzzles47%4%
LMArena Hard Prompts15261448
Mystery Game Puzzles58%5%
DTBench97.6%82.1%
LMCA65.5%28.3%
Epoch Capabilities Index164.7142.38
ARC-AGI-290%—
Kagi LLM Benchmark—72.5%
ARC-AGI-197.5%—
CritPt31.1%—
EBR-Bench57.1%—

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
FrontierMath (Tiers 1-3)90.2%18.9%
OTIS Mock AIME 2024-2025100%73.3%
LMArena Math15251446
FrontierMath Tier 487.8%—
ProofBench100%—
MATH Level 5—97.1%
FrontierMath Erdős0%—

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
SimpleQA Verified70.8%48.7%
LMArena Expert15351455
GPQA Diamond—72.6%
Humanity's Last Exam46.5%—

Multimodal Not comparable

Claude Fable 5.1: 53.9 (#4), Qwen3 Max: —

Multimodal benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
LMArena Vision1318—
Blueprint-Bench 241.9%—
Furniture Assembly70%—
LMArena Document1513—

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
LMArena Non-English15071429
LMArena Chinese15861478
LMArena French15251449
LMArena German15001463
LMArena Japanese15431397
LMArena Korean15341399
LMArena Russian15211428
LMArena Spanish15161462

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
LMArena Instruction Following15171419

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
LMArena Longer Query15221438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1Qwen3 Max
LMArena Text15101439
LMArena Creative Writing15071402
LMArena Multi-Turn14921446
EQ-Bench Creative Writing2162—

Frequently asked questions

Is Claude Fable 5.1 better than Qwen3 Max?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 43.7 on the Noometry Index. Qwen3 Max costs 8.3× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5.1 or Qwen3 Max?

Qwen3 Max is cheaper. It lists at $1.20 per million input tokens and $6 per million output tokens; Claude Fable 5.1 lists at $10 and $50.

Is Claude Fable 5.1 or Qwen3 Max better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 43.0 in the Noometry coding category.

Which has the bigger context window?

Claude Fable 5.1 does, with 1M tokens against 262K.

How many benchmarks do Claude Fable 5.1 and Qwen3 Max share?

28 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper