Model comparison

Claude 3 Opus vs Gemma 3 27B

Gemma 3 27B is the stronger model overall, scoring 30.8 to 29.5 on the Noometry Index.

Last verified . 33 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

Gemma 3 27B Google

30.8

Rank #284 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Claude 3 Opus scores higher in 2 categories and Gemma 3 27B in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 3 27B leads 25.9 to 14.8.
  • The biggest single-benchmark swing is MATH Level 5: 37.5% for Claude 3 Opus and 74% for Gemma 3 27B.
  • Gemma 3 27B has downloadable open weights; the other is API-only.

Side by side

Claude 3 Opus and Gemma 3 27B specifications
Claude 3 OpusGemma 3 27B
ProviderAnthropicGoogle
Noometry Index29.530.8
Released2024-02-292025-03-11
WeightsProprietaryOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.08
Output $ / M tokens—$0.16
Results tracked4643

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3 Opus leads

Claude 3 Opus: 32.9 (#267), Gemma 3 27B: 22.5 (#334)

Coding benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
LiveBench Coding38.6%39.9%
LMArena Coding12641322
Aider Polyglot—4.9%
SciCode—21.2%
WeirdML19.2%—
BigCodeBench Instruct45.5%—
BigCodeBench Complete57.4%—
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use Too close to call

Claude 3 Opus: 24.6 (#116), Gemma 3 27B: 25.1 (#110)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
Berkeley Function Calling Leaderboard—29.5%
Cybench10%—
METR Time Horizons29.5%—

Reasoning Gemma 3 27B leads

Claude 3 Opus: 14.6 (#324), Gemma 3 27B: 16.7 (#301)

Reasoning benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
Chess Puzzles5%0%
LiveBench Reasoning40.6%43.8%
LMArena Hard Prompts12451340
DTBench61.6%52.5%
LiveBench Data Analysis57.9%51.5%
LMCA17%12.3%
Epoch Capabilities Index126.91130.04
LiveBench49.2%50%
SimpleBench23.5%—
Kagi LLM Benchmark—40.4%
CritPt—0%
EnigmaEval0.8%—
ForecastBench58.4—
WinoGrande88.5%—

Math Gemma 3 27B leads

Claude 3 Opus: 14.8 (#299), Gemma 3 27B: 25.9 (#265)

Math benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
OTIS Mock AIME 2024-20254.7%22.5%
LiveBench Math43.6%55.4%
LMArena Math12731312
MATH Level 537.5%74%

Knowledge Gemma 3 27B leads

Claude 3 Opus: 24.5 (#267), Gemma 3 27B: 25.5 (#261)

Knowledge benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
GPQA Diamond47.2%47.7%
Confabulations22.7%40.3%
LMArena Expert12231304
SimpleQA Verified12.6%—
Vectara Hallucination Rate—7.4%
MMLU84.6%—

Multimodal Gemma 3 27B leads

Claude 3 Opus: 27.1 (#116), Gemma 3 27B: 32.6 (#100)

Multimodal benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
LMArena Vision10231164
GeoBench—52%

Multilingual Gemma 3 27B leads

Claude 3 Opus: 41.4 (#207), Gemma 3 27B: 46.9 (#155)

Multilingual benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
LMArena Non-English12581334
LMArena Chinese12481346
LMArena French12751368
LMArena German12581362
LMArena Japanese12041287
LMArena Korean11871308
LMArena Russian12801349
LMArena Spanish12461349

Instruction Following Gemma 3 27B leads

Claude 3 Opus: 64.1 (#228), Gemma 3 27B: 70.6 (#160)

Instruction Following benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
LiveBench Instruction Following63.9%74.9%
LMArena Instruction Following12481321

Long Context Claude 3 Opus leads

Claude 3 Opus: 38.2 (#202), Gemma 3 27B: 27.6 (#293)

Long Context benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
LMArena Longer Query12591333
Fiction.LiveBench—33.3%

Writing & Preference Gemma 3 27B leads

Claude 3 Opus: 47.2 (#213), Gemma 3 27B: 52.5 (#168)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusGemma 3 27B
LMArena Text12621358
LMArena Creative Writing12351346
LMArena Multi-Turn12751345
LiveBench Language50.4%34.6%
Short-Story Creative Writing—79.9%
EQ-Bench Creative Writing—1266

Frequently asked questions

Is Claude 3 Opus better than Gemma 3 27B?

Gemma 3 27B is the stronger model overall, scoring 30.8 to 29.5 on the Noometry Index.

Is Claude 3 Opus or Gemma 3 27B better for coding?

Claude 3 Opus scores higher on coding benchmarks: 32.9 versus 22.5 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and Gemma 3 27B share?

33 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and Gemma 3 27B has 43.

Related comparisons

Go deeper