Model comparison

Gemma 2B vs Llama 13b

Gemma 2B is the stronger model overall, scoring 29.6 to 24.4 on the Noometry Index.

Last verified . 18 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemma 2B scores higher in 6 categories and Llama 13b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Gemma 2B leads 48.5 to 36.7.

Side by side

Gemma 2B and Llama 13b specifications
Gemma 2BLlama 13b
ProviderGoogleMeta
Noometry Index29.624.4
Released2024-02-212023-02-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2321

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2B leads

Gemma 2B: 29.4 (#305), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkGemma 2BLlama 13b
LMArena Coding1010683
HumanEval+20.7%—
MBPP+34.1%—

Reasoning Gemma 2B leads

Gemma 2B: 18.8 (#275), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkGemma 2BLlama 13b
LMArena Hard Prompts989728
BIG-Bench Hard35.2%37.9%
Epoch Capabilities Index94.2100.58
HellaSwag71.4%79.2%
PIQA77.3%80.1%
WinoGrande65.4%73%
LAMBADA—75.2%

Math Gemma 2B leads

Gemma 2B: 30.0 (#239), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkGemma 2BLlama 13b
LMArena Math1009838
GSM8K17.7%20.6%

Knowledge Not comparable

Gemma 2B: —, Llama 13b: —

Knowledge benchmarks
BenchmarkGemma 2BLlama 13b
ARC (AI2) Challenge42.1%52.7%
BoolQ69.4%78.7%
MMLU42.3%47.7%
TriviaQA53.2%77.9%
OpenBookQA—56.4%

Multimodal Not comparable

Gemma 2B: —, Llama 13b: —

Multimodal benchmarks
BenchmarkGemma 2BLlama 13b
ScienceQA—43.3%

Multilingual Gemma 2B leads

Gemma 2B: 23.0 (#294), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkGemma 2BLlama 13b
LMArena Non-English958819
LMArena Chinese986—
LMArena Russian937—

Instruction Following Gemma 2B leads

Gemma 2B: 48.5 (#302), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkGemma 2BLlama 13b
LMArena Instruction Following970781

Long Context Not comparable

Gemma 2B: 29.9 (#291), Llama 13b: —

Long Context benchmarks
BenchmarkGemma 2BLlama 13b
LMArena Longer Query981—

Writing & Preference Gemma 2B leads

Gemma 2B: 24.0 (#308), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkGemma 2BLlama 13b
LMArena Text1002834
LMArena Creative Writing987794
LMArena Multi-Turn945753

Frequently asked questions

Is Gemma 2B better than Llama 13b?

Gemma 2B is the stronger model overall, scoring 29.6 to 24.4 on the Noometry Index.

Is Gemma 2B or Llama 13b better for coding?

Gemma 2B scores higher on coding benchmarks: 29.4 versus 21.4 in the Noometry coding category.

How many benchmarks do Gemma 2B and Llama 13b share?

18 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper