Model comparison

Claude Fable 5.1 vs Gemini 2.0 Flash (Feb 2025)

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 35.1 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude Fable 5.1 scores higher in 10 categories and Gemini 2.0 Flash (Feb 2025) in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Fable 5.1 leads 76.7 to 15.2.
  • The biggest single-benchmark swing is ARC-AGI-2: 90% for Claude Fable 5.1 and 1.3% for Gemini 2.0 Flash (Feb 2025).

Side by side

Claude Fable 5.1 and Gemini 2.0 Flash (Feb 2025) specifications
Claude Fable 5.1Gemini 2.0 Flash (Feb 2025)
ProviderAnthropicGoogle
Noometry Index69.035.1
Released2026-09-012024-12-06
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$10—
Output $ / M tokens$50—
Results tracked5254

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), Gemini 2.0 Flash (Feb 2025): 28.4 (#315)

Coding benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
WeirdML92.9%25.8%
LMArena Coding15281350
FrontierCode50.9%—
SWE-bench Verified (bash only)—13.5%
Aider Polyglot—38.2%
CursorBench51.8%—
LMArena WebDev1744—
FrontierSWE56.3%—
SciCode63.1%—
GSO88.2%—
BigCodeBench Instruct—45.9%
LiveBench Coding—63.4%
MirrorCode73.3%—
BigCodeBench Complete—59.9%
CadEval—30%
ALE-Bench2,143—

Agentic & Tool Use Claude Fable 5.1 leads

Claude Fable 5.1: 50.7 (#5), Gemini 2.0 Flash (Feb 2025): 28.1 (#92)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
APEX-Agents68.6%—
Remote Labor Index17.9%—
TheAgentCompany—11.4%
GDP.pdf29.6%—
Vending-Bench 25,422—

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), Gemini 2.0 Flash (Feb 2025): 15.2 (#318)

Reasoning benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
ARC-AGI-290%1.3%
LMArena Hard Prompts15261346
DTBench97.6%63.2%
Epoch Capabilities Index164.7135.36
SimpleBench—31.1%
Kagi LLM Benchmark—37.8%
NYT Connections (extended)90%—
ARC-AGI-197.5%—
CritPt31.1%—
Chess Puzzles47%—
EnigmaEval—1.1%
EBR-Bench57.1%—
LiveBench Reasoning—78.2%
Mystery Game Puzzles58%—
LiveBench Data Analysis—69.4%
LMCA65.5%—
LiveBench—66.9%

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), Gemini 2.0 Flash (Feb 2025): 37.9 (#146)

Math benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
OTIS Mock AIME 2024-2025100%57.8%
LMArena Math15251352
FrontierMath (Tiers 1-3)90.2%—
FrontierMath Tier 487.8%—
ProofBench100%—
Omni-MATH—45.9%
LiveBench Math—75.8%
MATH Level 5—82.2%
FrontierMath (Feb 2025 set)—1.7%
FrontierMath Erdős0%—

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), Gemini 2.0 Flash (Feb 2025): 32.0 (#213)

Knowledge benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
Humanity's Last Exam46.5%6.6%
LMArena Expert15351339
GPQA Diamond—64.1%
SimpleQA Verified70.8%—
MMLU-Pro—73.7%
Confabulations—12.4%
GPQA (HELM)—55.6%
MMLU—79.7%

Multimodal Claude Fable 5.1 leads

Claude Fable 5.1: 53.9 (#4), Gemini 2.0 Flash (Feb 2025): 36.5 (#79)

Multimodal benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
LMArena Vision13181158
GeoBench—77%
Blueprint-Bench 241.9%—
Furniture Assembly70%—
LMArena Document1513—

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), Gemini 2.0 Flash (Feb 2025): 47.4 (#149)

Multilingual benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
LMArena Non-English15071342
LMArena Chinese15861373
LMArena French15251391
LMArena German15001353
LMArena Japanese15431294
LMArena Korean15341313
LMArena Russian15211351
LMArena Spanish15161363

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), Gemini 2.0 Flash (Feb 2025): 74.4 (#97)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
LMArena Instruction Following15171336
LiveBench Instruction Following—85.8%
IFEval—84.1%

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), Gemini 2.0 Flash (Feb 2025): 38.1 (#203)

Long Context benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
LMArena Longer Query15221344
Fiction.LiveBench—61.1%

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), Gemini 2.0 Flash (Feb 2025): 49.5 (#190)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1Gemini 2.0 Flash (Feb 2025)
LMArena Text15101354
LMArena Creative Writing15071340
EQ-Bench Creative Writing21621128
LMArena Multi-Turn14921350
Short-Story Creative Writing—73.8%
WildBench—80%
LiveBench Language—51.3%

Frequently asked questions

Is Claude Fable 5.1 better than Gemini 2.0 Flash (Feb 2025)?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 35.1 on the Noometry Index.

Is Claude Fable 5.1 or Gemini 2.0 Flash (Feb 2025) better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 28.4 in the Noometry coding category.

How many benchmarks do Claude Fable 5.1 and Gemini 2.0 Flash (Feb 2025) share?

25 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and Gemini 2.0 Flash (Feb 2025) has 54.

Related comparisons

Go deeper