Model comparison

Claude 2.1 vs Claude Fable 5.1

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 25.2 on the Noometry Index.

Last verified . 4 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Claude Fable 5.1 in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Fable 5.1 leads 89.6 to 10.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.9% for Claude 2.1 and 100% for Claude Fable 5.1.

Side by side

Claude 2.1 and Claude Fable 5.1 specifications
Claude 2.1Claude Fable 5.1
ProviderAnthropicAnthropic
Noometry Index25.269.0
Released2023-11-212026-09-01
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$10
Output $ / M tokens—$50
Results tracked752

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude 2.1: 26.2 (#327), Claude Fable 5.1: 74.7 (#1)

Coding benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
WeirdML7.1%92.9%
FrontierCode—50.9%
CursorBench—51.8%
LMArena WebDev—1744
FrontierSWE—56.3%
SciCode—63.1%
GSO—88.2%
LMArena Coding—1528
MirrorCode—73.3%
ALE-Bench—2,143

Agentic & Tool Use Not comparable

Claude 2.1: —, Claude Fable 5.1: 50.7 (#5)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
APEX-Agents—68.6%
Remote Labor Index—17.9%
GDP.pdf—29.6%
Vending-Bench 2—5,422

Reasoning Claude Fable 5.1 leads

Claude 2.1: 21.4 (#221), Claude Fable 5.1: 76.7 (#7)

Reasoning benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
DTBench51%97.6%
Epoch Capabilities Index119.27164.7
ARC-AGI-2—90%
NYT Connections (extended)—90%
ARC-AGI-1—97.5%
CritPt—31.1%
Chess Puzzles—47%
EBR-Bench—57.1%
LMArena Hard Prompts—1526
Mystery Game Puzzles—58%
LMCA—65.5%
ForecastBench54.2—

Math Claude Fable 5.1 leads

Claude 2.1: 10.2 (#315), Claude Fable 5.1: 89.6 (#4)

Math benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
OTIS Mock AIME 2024-20251.9%100%
FrontierMath (Tiers 1-3)—90.2%
FrontierMath Tier 4—87.8%
ProofBench—100%
LMArena Math—1525
FrontierMath Erdős—0%

Knowledge Claude Fable 5.1 leads

Claude 2.1: 15.4 (#292), Claude Fable 5.1: 69.6 (#6)

Knowledge benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
GPQA Diamond33%—
Humanity's Last Exam—46.5%
SimpleQA Verified—70.8%
LMArena Expert—1535
MMLU73.5%—

Multimodal Not comparable

Claude 2.1: —, Claude Fable 5.1: 53.9 (#4)

Multimodal benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
LMArena Vision—1318
Blueprint-Bench 2—41.9%
Furniture Assembly—70%
LMArena Document—1513

Multilingual Not comparable

Claude 2.1: —, Claude Fable 5.1: 59.1 (#3)

Multilingual benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
LMArena Non-English—1507
LMArena Chinese—1586
LMArena French—1525
LMArena German—1500
LMArena Japanese—1543
LMArena Korean—1534
LMArena Russian—1521
LMArena Spanish—1516

Instruction Following Not comparable

Claude 2.1: —, Claude Fable 5.1: 79.2 (#6)

Instruction Following benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
LMArena Instruction Following—1517

Long Context Not comparable

Claude 2.1: —, Claude Fable 5.1: 46.7 (#20)

Long Context benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
LMArena Longer Query—1522

Writing & Preference Not comparable

Claude 2.1: —, Claude Fable 5.1: 79.2 (#2)

Writing & Preference benchmarks
BenchmarkClaude 2.1Claude Fable 5.1
LMArena Text—1510
LMArena Creative Writing—1507
EQ-Bench Creative Writing—2162
LMArena Multi-Turn—1492

Frequently asked questions

Is Claude 2.1 better than Claude Fable 5.1?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 25.2 on the Noometry Index.

Is Claude 2.1 or Claude Fable 5.1 better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Claude Fable 5.1 share?

4 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Claude Fable 5.1 has 52.

Related comparisons

Go deeper