Model comparison

Gemma 3 4B vs Llama 3.2 3B

Gemma 3 4B and Llama 3.2 3B score almost the same on the Noometry Index (28.1 vs 28.9), so choose on price, context window or the category you care about most.

Last verified . 14 shared benchmarks.

Gemma 3 4B Google

28.1

Rank #326 Confirmed

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 3 4B scores higher in 6 categories and Llama 3.2 3B in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 3.2 3B leads 29.7 to 11.8.
  • Gemma 3 4B is cheaper at $0.04 / $0.08 per million input/output tokens, against $0.05 / $0.33 for Llama 3.2 3B.

Side by side

Gemma 3 4B and Llama 3.2 3B specifications
Gemma 3 4BLlama 3.2 3B
ProviderGoogleMeta
Noometry Index28.128.9
Released2025-03-122024-09-24
WeightsOpenOpen
Context window131K131K
Max output4K118K
Input $ / M tokens$0.04$0.05
Output $ / M tokens$0.08$0.33
Results tracked2218

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3 4B leads

Gemma 3 4B: 35.9 (#215), Llama 3.2 3B: 27.6 (#319)

Coding benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Coding12301098
BigCodeBench Instruct—23.4%
BigCodeBench Complete—28.3%

Agentic & Tool Use Too close to call

Gemma 3 4B: 20.9 (#142), Llama 3.2 3B: 20.1 (#143)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
Berkeley Function Calling Leaderboard19.6%21.9%
BALROG—10.1%

Reasoning Llama 3.2 3B leads

Gemma 3 4B: 13.2 (#335), Llama 3.2 3B: 21.0 (#228)

Reasoning benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Hard Prompts12531095
Kagi LLM Benchmark25.2%—
Chess Puzzles0%—
DTBench50.9%—
LMCA2.8%—
Epoch Capabilities Index116.02—

Math Llama 3.2 3B leads

Gemma 3 4B: 16.8 (#292), Llama 3.2 3B: 32.4 (#214)

Math benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Math12391126
OTIS Mock AIME 2024-20257.5%—

Knowledge Llama 3.2 3B leads

Gemma 3 4B: 11.8 (#299), Llama 3.2 3B: 29.7 (#235)

Knowledge benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Expert12231090
GPQA Diamond23.2%—
Vectara Hallucination Rate6.4%—

Multilingual Gemma 3 4B leads

Gemma 3 4B: 42.5 (#194), Llama 3.2 3B: 26.2 (#281)

Multilingual benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Non-English12731019
LMArena German12811056
LMArena Russian1294949
LMArena Chinese—1017

Instruction Following Gemma 3 4B leads

Gemma 3 4B: 65.2 (#225), Llama 3.2 3B: 56.0 (#275)

Instruction Following benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Instruction Following12391089

Long Context Gemma 3 4B leads

Gemma 3 4B: 38.7 (#194), Llama 3.2 3B: 33.4 (#261)

Long Context benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Longer Query12731100

Writing & Preference Gemma 3 4B leads

Gemma 3 4B: 42.0 (#239), Llama 3.2 3B: 24.7 (#307)

Writing & Preference benchmarks
BenchmarkGemma 3 4BLlama 3.2 3B
LMArena Text12911110
LMArena Creative Writing12711094
EQ-Bench Creative Writing1068595
LMArena Multi-Turn12551105

Frequently asked questions

Is Gemma 3 4B better than Llama 3.2 3B?

Gemma 3 4B and Llama 3.2 3B score almost the same on the Noometry Index (28.1 vs 28.9), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 3 4B or Llama 3.2 3B?

Gemma 3 4B is cheaper. It lists at $0.04 per million input tokens and $0.08 per million output tokens; Llama 3.2 3B lists at $0.05 and $0.33.

Is Gemma 3 4B or Llama 3.2 3B better for coding?

Gemma 3 4B scores higher on coding benchmarks: 35.9 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Gemma 3 4B and Llama 3.2 3B share?

14 benchmarks have published results for both models. Gemma 3 4B has 22 scored results on Noometry and Llama 3.2 3B has 18.

Related comparisons

Go deeper