Model comparison

Claude 3.7 Sonnet vs Gemini 2.5 Flash

Claude 3.7 Sonnet and Gemini 2.5 Flash score almost the same on the Noometry Index (39.5 vs 39.3), so choose on price, context window or the category you care about most.

Last verified . 42 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

Summary

  • They share 42 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 6 categories and Gemini 2.5 Flash in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Gemini 2.5 Flash leads 52.3 to 44.1.
  • The biggest single-benchmark swing is SWE-bench Verified (bash only): 52.8% for Claude 3.7 Sonnet and 28.7% for Gemini 2.5 Flash.

Side by side

Claude 3.7 Sonnet and Gemini 2.5 Flash specifications
Claude 3.7 SonnetGemini 2.5 Flash
ProviderAnthropicGoogle
Noometry Index39.539.3
Released2025-02-242025-04-17
WeightsProprietaryProprietary
Context window—1.05M
Max output—66K
Input $ / M tokens—$0.30
Output $ / M tokens—$2.50
Results tracked5854

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 40.6 (#136), Gemini 2.5 Flash: 35.8 (#220)

Coding benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
SWE-bench Verified (bash only)52.8%28.7%
Aider Polyglot64.9%55.1%
LMArena Coding13611424
SWE-bench Verified61%—
GSO3.8%—
WeirdML—41.9%
LiveBench Coding74.5%—
CadEval54%—
ALE-Bench—661.88

Agentic & Tool Use Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 34.1 (#50), Gemini 2.5 Flash: 30.8 (#74)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
TheAgentCompany30.9%41.1%
Terminal-Bench—17.1%
Berkeley Function Calling Leaderboard—56.2%
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—
BALROG—33.5%
METR Time Horizons60%—
Vending-Bench 2—548.84

Reasoning Too close to call

Claude 3.7 Sonnet: 18.6 (#277), Gemini 2.5 Flash: 18.1 (#286)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
ARC-AGI-20.9%2.5%
SimpleBench46.4%41.2%
ARC-AGI-128.6%33.3%
EnigmaEval4.2%2.7%
LMArena Hard Prompts13331422
Epoch Capabilities Index141.16143.03
ForecastBench61.860.6
Kagi LLM Benchmark—56.8%
CritPt—1.1%
LiveBench Reasoning87.8%—
DTBench—76.5%
LiveBench Data Analysis74%—
LMCA—27.5%
LiveBench76.1%—

Math Gemini 2.5 Flash leads

Claude 3.7 Sonnet: 37.5 (#153), Gemini 2.5 Flash: 39.9 (#98)

Math benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
OTIS Mock AIME 2024-202557.8%73.1%
Omni-MATH33%38.5%
LMArena Math13371415
FrontierMath (Feb 2025 set)4.1%4.8%
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath Tier 4 (v1)—4.2%

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), Gemini 2.5 Flash: 36.4 (#168)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
Humanity's Last Exam8%12.1%
MMLU-Pro78.4%63.9%
Confabulations14.7%16.8%
GPQA (HELM)60.8%39%
LMArena Expert13211426
GPQA Diamond79.7%—
Vectara Hallucination Rate—7.8%

Multimodal Gemini 2.5 Flash leads

Claude 3.7 Sonnet: 33.7 (#95), Gemini 2.5 Flash: 41.8 (#32)

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
LMArena Vision11691253
GeoBench68%76%
VPCT39%46.2%
SpatialViz-Bench33.9%36.9%

Multilingual Gemini 2.5 Flash leads

Claude 3.7 Sonnet: 44.1 (#179), Gemini 2.5 Flash: 52.3 (#88)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
LMArena Non-English12961409
LMArena Chinese12991450
LMArena French13031433
LMArena German13011418
LMArena Japanese12671405
LMArena Korean12491385
LMArena Russian13111415
LMArena Spanish12981421

Instruction Following Gemini 2.5 Flash leads

Claude 3.7 Sonnet: 72.9 (#125), Gemini 2.5 Flash: 75.7 (#54)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
IFEval83.4%89.8%
LMArena Instruction Following13521405
LiveBench Instruction Following81.3%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Gemini 2.5 Flash: 47.5 (#17)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
Fiction.LiveBench83.3%77.8%
LMArena Longer Query13731419

Writing & Preference Too close to call

Claude 3.7 Sonnet: 54.4 (#150), Gemini 2.5 Flash: 53.8 (#157)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetGemini 2.5 Flash
LMArena Text13141417
LMArena Creative Writing13321400
Short-Story Creative Writing81.1%76.5%
EQ-Bench Creative Writing14121137
WildBench81.4%81.7%
LMArena Multi-Turn13391408
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Gemini 2.5 Flash?

Claude 3.7 Sonnet and Gemini 2.5 Flash score almost the same on the Noometry Index (39.5 vs 39.3), so choose on price, context window or the category you care about most.

Is Claude 3.7 Sonnet or Gemini 2.5 Flash better for coding?

Claude 3.7 Sonnet scores higher on coding benchmarks: 40.6 versus 35.8 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Gemini 2.5 Flash share?

42 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Gemini 2.5 Flash has 54.

Related comparisons

Go deeper