Model comparison

Claude Fable 5.1 vs DeepSeek V4 Pro

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 54.3 on the Noometry Index. DeepSeek V4 Pro costs 20× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Last verified . 39 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Claude Fable 5.1 scores higher in 9 categories and DeepSeek V4 Pro in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Fable 5.1 leads 89.6 to 64.8.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 87.8% for Claude Fable 5.1 and 26.8% for DeepSeek V4 Pro.
  • DeepSeek V4 Pro is cheaper at $0.66 / $1.98 per million input/output tokens, against $10 / $50 for Claude Fable 5.1.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

Claude Fable 5.1 and DeepSeek V4 Pro specifications
Claude Fable 5.1DeepSeek V4 Pro
ProviderAnthropicDeepSeek
Noometry Index69.054.3
Released2026-09-012026-04-24
WeightsProprietaryOpen
Context window1M1M
Max output128K393K
Input $ / M tokens$10$0.66
Output $ / M tokens$50$1.98
Results tracked5248

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), DeepSeek V4 Pro: 52.4 (#34)

Coding benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
FrontierCode50.9%28.6%
LMArena WebDev17441582
SciCode63.1%51%
WeirdML92.9%66.2%
LMArena Coding15281470
ALE-Bench2,1431,403
SWE-bench Verified—77.6%
CursorBench51.8%—
FrontierSWE56.3%—
GSO88.2%—
MirrorCode73.3%—

Agentic & Tool Use Claude Fable 5.1 leads

Claude Fable 5.1: 50.7 (#5), DeepSeek V4 Pro: 32.8 (#58)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
APEX-Agents68.6%47.3%
Vending-Bench 25,4223,285
Remote Labor Index17.9%—
GDP.pdf29.6%—

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), DeepSeek V4 Pro: 56.5 (#24)

Reasoning benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
ARC-AGI-290%61.3%
NYT Connections (extended)90%91.3%
ARC-AGI-197.5%90.5%
CritPt31.1%18%
Chess Puzzles47%47%
LMArena Hard Prompts15261461
Mystery Game Puzzles58%43%
DTBench97.6%93.9%
LMCA65.5%45.5%
Epoch Capabilities Index164.7155.31
Kagi LLM Benchmark—53.5%
EBR-Bench57.1%—
Surface Evolver Bench—40%
ForecastBench—56.1

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), DeepSeek V4 Pro: 64.8 (#30)

Math benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
FrontierMath (Tiers 1-3)90.2%64.6%
FrontierMath Tier 487.8%26.8%
OTIS Mock AIME 2024-2025100%98.6%
ProofBench100%50%
LMArena Math15251455
MathArena Final-Answer Competitions—76.6%
FrontierMath Erdős0%—

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), DeepSeek V4 Pro: 59.5 (#31)

Knowledge benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
SimpleQA Verified70.8%52.9%
LMArena Expert15351464
GPQA Diamond—91.7%
Humanity's Last Exam46.5%—
Vectara Hallucination Rate—8.6%

Multimodal Not comparable

Claude Fable 5.1: 53.9 (#4), DeepSeek V4 Pro: —

Multimodal benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
LMArena Vision1318—
Blueprint-Bench 241.9%—
Furniture Assembly70%—
LMArena Document1513—

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), DeepSeek V4 Pro: 54.4 (#45)

Multilingual benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
LMArena Non-English15071439
LMArena Chinese15861486
LMArena French15251472
LMArena German15001458
LMArena Japanese15431445
LMArena Korean15341447
LMArena Russian15211453
LMArena Spanish15161458

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), DeepSeek V4 Pro: 76.1 (#47)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
LMArena Instruction Following15171448

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), DeepSeek V4 Pro: 45.0 (#51)

Long Context benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
LMArena Longer Query15221458
CL-bench Life—13.5%

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), DeepSeek V4 Pro: 65.5 (#46)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1DeepSeek V4 Pro
LMArena Text15101451
LMArena Creative Writing15071446
EQ-Bench Creative Writing21621553
LMArena Multi-Turn14921467
EQ-Bench 4—1166

Frequently asked questions

Is Claude Fable 5.1 better than DeepSeek V4 Pro?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 54.3 on the Noometry Index. DeepSeek V4 Pro costs 20× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5.1 or DeepSeek V4 Pro?

DeepSeek V4 Pro is cheaper. It lists at $0.66 per million input tokens and $1.98 per million output tokens; Claude Fable 5.1 lists at $10 and $50.

Is Claude Fable 5.1 or DeepSeek V4 Pro better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 52.4 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Claude Fable 5.1 and DeepSeek V4 Pro share?

39 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and DeepSeek V4 Pro has 48.

Related comparisons

Go deeper