Model comparison

Claude Fable 5.1 vs Gemini 3 Flash Preview

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 52.3 on the Noometry Index. Gemini 3 Flash Preview costs 18× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Last verified . 39 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Gemini 3 Flash Preview Google

52.3

Rank #40 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Claude Fable 5.1 scores higher in 10 categories and Gemini 3 Flash Preview in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Fable 5.1 leads 89.6 to 51.7.
  • The biggest single-benchmark swing is ProofBench: 100% for Claude Fable 5.1 and 15% for Gemini 3 Flash Preview.
  • Gemini 3 Flash Preview is cheaper at $0.50 / $3 per million input/output tokens, against $10 / $50 for Claude Fable 5.1.
  • Gemini 3 Flash Preview accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Fable 5.1 and Gemini 3 Flash Preview specifications
Claude Fable 5.1Gemini 3 Flash Preview
ProviderAnthropicGoogle
Noometry Index69.052.3
Released2026-09-012025-12-17
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K66K
Input $ / M tokens$10$0.50
Output $ / M tokens$50$3
Results tracked5259

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), Gemini 3 Flash Preview: 50.9 (#42)

Coding benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
LMArena WebDev17441439
GSO88.2%9.8%
WeirdML92.9%61.6%
LMArena Coding15281460
ALE-Bench2,1431,367
SWE-bench Verified—75.4%
FrontierCode50.9%—
SWE-bench Verified (bash only)—75.8%
CursorBench51.8%—
SWE-bench Multilingual—72.7%
FrontierSWE56.3%—
SciCode63.1%—
MirrorCode73.3%—

Agentic & Tool Use Claude Fable 5.1 leads

Claude Fable 5.1: 50.7 (#5), Gemini 3 Flash Preview: 38.7 (#29)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
GDP.pdf29.6%10%
Vending-Bench 25,4223,635
Terminal-Bench—64.3%
APEX-Agents68.6%—
Remote Labor Index17.9%—
τ²-bench Airline—82.5%
τ²-bench Banking—27.3%
τ²-bench Retail—76.8%
τ²-bench Telecom—91.2%
DeepResearch Bench—49.8%
BALROG—48.1%
LMArena Search—1198

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), Gemini 3 Flash Preview: 49.2 (#37)

Reasoning benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
ARC-AGI-290%33.6%
NYT Connections (extended)90%83.1%
ARC-AGI-197.5%84.7%
Chess Puzzles47%40%
LMArena Hard Prompts15261465
Mystery Game Puzzles58%26%
DTBench97.6%89.1%
LMCA65.5%43.1%
Epoch Capabilities Index164.7151.8
SimpleBench—61.1%
CritPt31.1%—
EBR-Bench57.1%—
ForecastBench—58.5

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), Gemini 3 Flash Preview: 51.7 (#55)

Math benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
FrontierMath (Tiers 1-3)90.2%51.2%
FrontierMath Tier 487.8%17.1%
OTIS Mock AIME 2024-2025100%95.6%
ProofBench100%15%
LMArena Math15251473
MathArena Final-Answer Competitions—67.6%
FrontierMath (Feb 2025 set)—35.6%
FrontierMath Erdős0%—
FrontierMath Tier 4 (v1)—4.2%

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), Gemini 3 Flash Preview: 58.8 (#33)

Knowledge benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
SimpleQA Verified70.8%66.8%
LMArena Expert15351462
GPQA Diamond—89.4%
Humanity's Last Exam46.5%—
Vectara Hallucination Rate—13.5%

Multimodal Claude Fable 5.1 leads

Claude Fable 5.1: 53.9 (#4), Gemini 3 Flash Preview: 45.5 (#16)

Multimodal benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
LMArena Vision13181285
Blueprint-Bench 241.9%0%
LMArena Document15131413
GeoBench—88%
VPCT—72.6%
Furniture Assembly70%—

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), Gemini 3 Flash Preview: 55.7 (#27)

Multilingual benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
LMArena Non-English15071458
LMArena Chinese15861511
LMArena French15251477
LMArena German15001497
LMArena Japanese15431489
LMArena Korean15341443
LMArena Russian15211480
LMArena Spanish15161469

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), Gemini 3 Flash Preview: 75.7 (#56)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
LMArena Instruction Following15171437

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), Gemini 3 Flash Preview: 44.4 (#67)

Long Context benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
LMArena Longer Query15221452

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), Gemini 3 Flash Preview: 65.5 (#45)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1Gemini 3 Flash Preview
LMArena Text15101466
LMArena Creative Writing15071457
LMArena Multi-Turn14921471
EQ-Bench Creative Writing2162—

Frequently asked questions

Is Claude Fable 5.1 better than Gemini 3 Flash Preview?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 52.3 on the Noometry Index. Gemini 3 Flash Preview costs 18× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5.1 or Gemini 3 Flash Preview?

Gemini 3 Flash Preview is cheaper. It lists at $0.50 per million input tokens and $3 per million output tokens; Claude Fable 5.1 lists at $10 and $50.

Is Claude Fable 5.1 or Gemini 3 Flash Preview better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 50.9 in the Noometry coding category.

Which has the bigger context window?

Gemini 3 Flash Preview does, with 1.05M tokens against 1M.

How many benchmarks do Claude Fable 5.1 and Gemini 3 Flash Preview share?

39 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and Gemini 3 Flash Preview has 59.

Related comparisons

Go deeper