Model comparison

Gemma 3 27B vs Llama 3-70B

Gemma 3 27B is the stronger model overall, scoring 30.8 to 28.8 on the Noometry Index.

Last verified . 23 shared benchmarks.

Gemma 3 27B Google

30.8

Rank #284 Confirmed

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Gemma 3 27B scores higher in 6 categories and Llama 3-70B in 3 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Llama 3-70B leads 35.8 to 22.5.
  • The biggest single-benchmark swing is MATH Level 5: 74% for Gemma 3 27B and 22.6% for Llama 3-70B.

Side by side

Gemma 3 27B and Llama 3-70B specifications
Gemma 3 27BLlama 3-70B
ProviderGoogleMeta
Noometry Index30.828.8
Released2025-03-112024-04-18
WeightsOpenOpen
Context window131K—
Max output8K—
Input $ / M tokens$0.08—
Output $ / M tokens$0.16—
Results tracked4331

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3-70B leads

Gemma 3 27B: 22.5 (#334), Llama 3-70B: 35.8 (#218)

Coding benchmarks
BenchmarkGemma 3 27BLlama 3-70B
LMArena Coding13221206
Aider Polyglot4.9%—
SciCode21.2%—
BigCodeBench Instruct—43.6%
LiveBench Coding39.9%—
BigCodeBench Complete—54.5%
HumanEval+—72%
MBPP+—69%

Agentic & Tool Use Gemma 3 27B leads

Gemma 3 27B: 25.1 (#110), Llama 3-70B: 21.1 (#139)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 27BLlama 3-70B
Berkeley Function Calling Leaderboard29.5%—
Cybench—5%

Reasoning Llama 3-70B leads

Gemma 3 27B: 16.7 (#301), Llama 3-70B: 18.0 (#288)

Reasoning benchmarks
BenchmarkGemma 3 27BLlama 3-70B
Kagi LLM Benchmark40.4%35.1%
LMArena Hard Prompts13401195
DTBench52.5%54.2%
Epoch Capabilities Index130.04122.93
CritPt0%—
Chess Puzzles0%—
LiveBench Reasoning43.8%—
LiveBench Data Analysis51.5%—
LMCA12.3%—
ForecastBench—57.1
LiveBench50%—
WinoGrande—83.5%

Math Gemma 3 27B leads

Gemma 3 27B: 25.9 (#265), Llama 3-70B: 12.8 (#305)

Math benchmarks
BenchmarkGemma 3 27BLlama 3-70B
OTIS Mock AIME 2024-202522.5%4.3%
LMArena Math13121218
MATH Level 574%22.6%
LiveBench Math55.4%—

Knowledge Gemma 3 27B leads

Gemma 3 27B: 25.5 (#261), Llama 3-70B: 20.8 (#277)

Knowledge benchmarks
BenchmarkGemma 3 27BLlama 3-70B
GPQA Diamond47.7%40.6%
LMArena Expert13041149
Confabulations40.3%—
Vectara Hallucination Rate7.4%—
MMLU—79.3%

Multimodal Not comparable

Gemma 3 27B: 32.6 (#100), Llama 3-70B: —

Multimodal benchmarks
BenchmarkGemma 3 27BLlama 3-70B
LMArena Vision1164—
GeoBench52%—

Multilingual Gemma 3 27B leads

Gemma 3 27B: 46.9 (#155), Llama 3-70B: 33.6 (#251)

Multilingual benchmarks
BenchmarkGemma 3 27BLlama 3-70B
LMArena Non-English13341142
LMArena Chinese13461114
LMArena French13681232
LMArena German13621169
LMArena Japanese12871017
LMArena Korean13081017
LMArena Russian13491159
LMArena Spanish13491241

Instruction Following Gemma 3 27B leads

Gemma 3 27B: 70.6 (#160), Llama 3-70B: 62.5 (#238)

Instruction Following benchmarks
BenchmarkGemma 3 27BLlama 3-70B
LMArena Instruction Following13211194
LiveBench Instruction Following74.9%—

Long Context Llama 3-70B leads

Gemma 3 27B: 27.6 (#293), Llama 3-70B: 35.6 (#240)

Long Context benchmarks
BenchmarkGemma 3 27BLlama 3-70B
LMArena Longer Query13331174
Fiction.LiveBench33.3%—

Writing & Preference Gemma 3 27B leads

Gemma 3 27B: 52.5 (#168), Llama 3-70B: 42.8 (#231)

Writing & Preference benchmarks
BenchmarkGemma 3 27BLlama 3-70B
LMArena Text13581221
LMArena Creative Writing13461210
LMArena Multi-Turn13451223
Short-Story Creative Writing79.9%—
EQ-Bench Creative Writing1266—
LiveBench Language34.6%—

Frequently asked questions

Is Gemma 3 27B better than Llama 3-70B?

Gemma 3 27B is the stronger model overall, scoring 30.8 to 28.8 on the Noometry Index.

Is Gemma 3 27B or Llama 3-70B better for coding?

Llama 3-70B scores higher on coding benchmarks: 35.8 versus 22.5 in the Noometry coding category.

How many benchmarks do Gemma 3 27B and Llama 3-70B share?

23 benchmarks have published results for both models. Gemma 3 27B has 43 scored results on Noometry and Llama 3-70B has 31.

Related comparisons

Go deeper