Model comparison

Gemma 3 4B vs Llama 3.2 90B

Gemma 3 4B and Llama 3.2 90B score almost the same on the Noometry Index (28.1 vs 27.5), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

Gemma 3 4B Google

28.1

Rank #326 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Gemma 3 4B scores higher in 1 category and Llama 3.2 90B in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama 3.2 90B leads 21.7 to 11.8.
  • The biggest single-benchmark swing is GPQA Diamond: 23.2% for Gemma 3 4B and 41% for Llama 3.2 90B.

Side by side

Gemma 3 4B and Llama 3.2 90B specifications
Gemma 3 4BLlama 3.2 90B
ProviderGoogleMeta
Noometry Index28.127.5
Released2025-03-122024-09-24
WeightsOpenOpen
Context window131K—
Max output4K—
Input $ / M tokens$0.04—
Output $ / M tokens$0.08—
Results tracked229

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Gemma 3 4B: 35.9 (#215), Llama 3.2 90B: —

Coding benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
LMArena Coding1230—

Agentic & Tool Use Llama 3.2 90B leads

Gemma 3 4B: 20.9 (#142), Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
Berkeley Function Calling Leaderboard19.6%—
BALROG—27.3%

Reasoning Llama 3.2 90B leads

Gemma 3 4B: 13.2 (#335), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
Epoch Capabilities Index116.02125.5
Kagi LLM Benchmark25.2%—
Chess Puzzles0%—
EnigmaEval—0.4%
LMArena Hard Prompts1253—
DTBench50.9%—
LMCA2.8%—

Math Gemma 3 4B leads

Gemma 3 4B: 16.8 (#292), Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
OTIS Mock AIME 2024-20257.5%2.6%
LMArena Math1239—
MATH Level 5—39.4%

Knowledge Llama 3.2 90B leads

Gemma 3 4B: 11.8 (#299), Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
GPQA Diamond23.2%41%
Vectara Hallucination Rate6.4%—
LMArena Expert1223—
MMLU—80.3%

Multimodal Not comparable

Gemma 3 4B: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
LMArena Vision—1000
GeoBench—52%

Multilingual Not comparable

Gemma 3 4B: 42.5 (#194), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
LMArena Non-English1273—
LMArena German1281—
LMArena Russian1294—

Instruction Following Not comparable

Gemma 3 4B: 65.2 (#225), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
LMArena Instruction Following1239—

Long Context Not comparable

Gemma 3 4B: 38.7 (#194), Llama 3.2 90B: —

Long Context benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
LMArena Longer Query1273—

Writing & Preference Not comparable

Gemma 3 4B: 42.0 (#239), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkGemma 3 4BLlama 3.2 90B
LMArena Text1291—
LMArena Creative Writing1271—
EQ-Bench Creative Writing1068—
LMArena Multi-Turn1255—

Frequently asked questions

Is Gemma 3 4B better than Llama 3.2 90B?

Gemma 3 4B and Llama 3.2 90B score almost the same on the Noometry Index (28.1 vs 27.5), so choose on price, context window or the category you care about most.

How many benchmarks do Gemma 3 4B and Llama 3.2 90B share?

3 benchmarks have published results for both models. Gemma 3 4B has 22 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper