Model comparison

Claude Fable 5.1 vs Gemini 1.5 Flash (May 2024)

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 33.2 on the Noometry Index.

Last verified . 22 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Claude Fable 5.1 scores higher in 10 categories and Gemini 1.5 Flash (May 2024) in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Fable 5.1 leads 89.6 to 22.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 100% for Claude Fable 5.1 and 16.3% for Gemini 1.5 Flash (May 2024).

Side by side

Claude Fable 5.1 and Gemini 1.5 Flash (May 2024) specifications
Claude Fable 5.1Gemini 1.5 Flash (May 2024)
ProviderAnthropicGoogle
Noometry Index69.033.2
Released2026-09-012024-05-14
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$10—
Output $ / M tokens$50—
Results tracked5242

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), Gemini 1.5 Flash (May 2024): 34.4 (#236)

Coding benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
WeirdML92.9%24.9%
LMArena Coding15281261
FrontierCode50.9%—
CursorBench51.8%—
LMArena WebDev1744—
FrontierSWE56.3%—
SciCode63.1%—
GSO88.2%—
BigCodeBench Instruct—43.5%
MirrorCode73.3%—
BigCodeBench Complete—55.1%
ALE-Bench2,143—
HumanEval+—75.6%
MBPP+—67.5%

Agentic & Tool Use Claude Fable 5.1 leads

Claude Fable 5.1: 50.7 (#5), Gemini 1.5 Flash (May 2024): 26.6 (#102)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
APEX-Agents68.6%—
Remote Labor Index17.9%—
BALROG—14.6%
GDP.pdf29.6%—
Vending-Bench 25,422—

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), Gemini 1.5 Flash (May 2024): 21.7 (#215)

Reasoning benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
LMArena Hard Prompts15261257
DTBench97.6%53.8%
Epoch Capabilities Index164.7129.36
ARC-AGI-290%—
NYT Connections (extended)90%—
ARC-AGI-197.5%—
CritPt31.1%—
Chess Puzzles47%—
EBR-Bench57.1%—
Mystery Game Puzzles58%—
LMCA65.5%—
ForecastBench—53.9
PIQA—87.5%

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), Gemini 1.5 Flash (May 2024): 22.1 (#281)

Math benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
OTIS Mock AIME 2024-2025100%16.3%
LMArena Math15251269
FrontierMath (Tiers 1-3)90.2%—
FrontierMath Tier 487.8%—
ProofBench100%—
Omni-MATH—30.4%
MATH Level 5—61.9%
FrontierMath (Feb 2025 set)—0%
FrontierMath Erdős0%—
GSM8K—82.4%

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), Gemini 1.5 Flash (May 2024): 26.2 (#260)

Knowledge benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
LMArena Expert15351233
GPQA Diamond—47.3%
Humanity's Last Exam46.5%—
SimpleQA Verified70.8%—
MMLU-Pro—67.8%
GPQA (HELM)—43.7%
BoolQ—85.8%
MMLU—77.9%

Multimodal Claude Fable 5.1 leads

Claude Fable 5.1: 53.9 (#4), Gemini 1.5 Flash (May 2024): 36.0 (#81)

Multimodal benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
LMArena Vision13181141
Video-MME—70.3%
GeoBench—76%
Blueprint-Bench 241.9%—
Furniture Assembly70%—
LMArena Document1513—

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), Gemini 1.5 Flash (May 2024): 42.9 (#189)

Multilingual benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
LMArena Non-English15071278
LMArena Chinese15861295
LMArena French15251258
LMArena German15001262
LMArena Japanese15431252
LMArena Korean15341221
LMArena Russian15211288
LMArena Spanish15161243

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), Gemini 1.5 Flash (May 2024): 66.8 (#205)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
LMArena Instruction Following15171258
IFEval—83.1%

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), Gemini 1.5 Flash (May 2024): 39.0 (#187)

Long Context benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
LMArena Longer Query15221284

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), Gemini 1.5 Flash (May 2024): 48.7 (#196)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1Gemini 1.5 Flash (May 2024)
LMArena Text15101287
LMArena Creative Writing15071285
LMArena Multi-Turn14921253
EQ-Bench Creative Writing2162—
WildBench—79.2%

Frequently asked questions

Is Claude Fable 5.1 better than Gemini 1.5 Flash (May 2024)?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 33.2 on the Noometry Index.

Is Claude Fable 5.1 or Gemini 1.5 Flash (May 2024) better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 34.4 in the Noometry coding category.

How many benchmarks do Claude Fable 5.1 and Gemini 1.5 Flash (May 2024) share?

22 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and Gemini 1.5 Flash (May 2024) has 42.

Related comparisons

Go deeper