Model comparison

Claude 3.5 Haiku vs Gemma 2 27B

Claude 3.5 Haiku and Gemma 2 27B score almost the same on the Noometry Index (29.2 vs 29.4), so choose on price, context window or the category you care about most.

Last verified . 33 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Gemma 2 27B Google

29.4

Rank #312 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 5 categories and Gemma 2 27B in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude 3.5 Haiku leads 14.7 to 10.7.
  • The biggest single-benchmark swing is MATH Level 5: 46.4% for Claude 3.5 Haiku and 27.9% for Gemma 2 27B.
  • Gemma 2 27B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Gemma 2 27B specifications
Claude 3.5 HaikuGemma 2 27B
ProviderAnthropicGoogle
Noometry Index29.229.4
Released2024-10-222024-06-24
WeightsProprietaryOpen
Context window—8K
Max output—2K
Input $ / M tokens—$0.65
Output $ / M tokens—$0.65
Results tracked4934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 27B leads

Claude 3.5 Haiku: 32.9 (#265), Gemma 2 27B: 34.1 (#246)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
BigCodeBench Instruct46.1%42.8%
LiveBench Coding51.4%36%
LMArena Coding12861211
BigCodeBench Complete59%52.5%
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Gemma 2 27B: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
BALROG19.3%—

Reasoning Claude 3.5 Haiku leads

Claude 3.5 Haiku: 17.7 (#290), Gemma 2 27B: 15.3 (#315)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
LiveBench Reasoning28.1%28.1%
LMArena Hard Prompts12511198
DTBench56.7%48%
LiveBench Data Analysis48.5%47.9%
Epoch Capabilities Index127.15122.08
LiveBench43.5%38.2%
CritPt0%—
LMCA—7.1%

Math Claude 3.5 Haiku leads

Claude 3.5 Haiku: 14.7 (#300), Gemma 2 27B: 10.7 (#311)

Math benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
OTIS Mock AIME 2024-20254.3%1.4%
LiveBench Math35.5%26.5%
LMArena Math12441212
MATH Level 546.4%27.9%
Omni-MATH22.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Too close to call

Claude 3.5 Haiku: 18.7 (#281), Gemma 2 27B: 19.0 (#280)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
GPQA Diamond38.1%36.5%
Confabulations36.7%27.1%
LMArena Expert12081172
MMLU74.3%75.7%
MMLU-Pro60.5%—
GPQA (HELM)36.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Gemma 2 27B: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
LMArena Vision1092—
GeoBench34%—

Multilingual Claude 3.5 Haiku leads

Claude 3.5 Haiku: 40.0 (#218), Gemma 2 27B: 38.6 (#226)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
LMArena Non-English12381217
LMArena Chinese12291221
LMArena French12641247
LMArena German12371209
LMArena Japanese11751175
LMArena Korean11731174
LMArena Russian12531234
LMArena Spanish12611228

Instruction Following Claude 3.5 Haiku leads

Claude 3.5 Haiku: 62.9 (#234), Gemma 2 27B: 60.5 (#249)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
LiveBench Instruction Following61.9%58.1%
LMArena Instruction Following12411206
IFEval79.2%—

Long Context Too close to call

Claude 3.5 Haiku: 38.3 (#200), Gemma 2 27B: 37.3 (#218)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
LMArena Longer Query12611231

Writing & Preference Gemma 2 27B leads

Claude 3.5 Haiku: 42.7 (#234), Gemma 2 27B: 44.2 (#225)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGemma 2 27B
LMArena Text12551231
LMArena Creative Writing12331241
LMArena Multi-Turn12651224
LiveBench Language35.4%32.6%
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—

Frequently asked questions

Is Claude 3.5 Haiku better than Gemma 2 27B?

Claude 3.5 Haiku and Gemma 2 27B score almost the same on the Noometry Index (29.2 vs 29.4), so choose on price, context window or the category you care about most.

Is Claude 3.5 Haiku or Gemma 2 27B better for coding?

Gemma 2 27B scores higher on coding benchmarks: 34.1 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Gemma 2 27B share?

33 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Gemma 2 27B has 34.

Related comparisons

Go deeper