Model comparison

Gemma 1.1 2b IT vs Llama 2-13B

Gemma 1.1 2b IT and Llama 2-13B score almost the same on the Noometry Index (29.3 vs 29.6), so choose on price, context window or the category you care about most.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Llama 2-13B Meta

29.6

Rank #309 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 1 category and Llama 2-13B in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemma 1.1 2b IT leads 19.1 to 12.8.

Side by side

Gemma 1.1 2b IT and Llama 2-13B specifications
Gemma 1.1 2b ITLlama 2-13B
ProviderGoogleMeta
Noometry Index29.329.6
Released—2023-07-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1632

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 1.1 2b IT: 30.1 (#299), Llama 2-13B: 30.9 (#291)

Coding benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Coding10341062
HumanEval+17.7%—
MBPP+23.3%—

Reasoning Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 19.1 (#270), Llama 2-13B: 12.8 (#337)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Hard Prompts10051051
Chess Puzzles—0%
DTBench—42.2%
BIG-Bench Hard—58.2%
Epoch Capabilities Index—106.17
HellaSwag—80.7%
LAMBADA—76.5%
PIQA—80.8%
WinoGrande—72.8%

Math Too close to call

Gemma 1.1 2b IT: 30.8 (#232), Llama 2-13B: 31.1 (#229)

Math benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Math10471065
GSM8K—36.9%

Knowledge Llama 2-13B leads

Gemma 1.1 2b IT: 26.5 (#258), Llama 2-13B: 28.1 (#249)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Expert9701030
ARC (AI2) Challenge—60.3%
BoolQ—82.4%
MMLU—55.6%
OpenBookQA—57%
TriviaQA—79.6%

Multimodal Not comparable

Gemma 1.1 2b IT: —, Llama 2-13B: —

Multimodal benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
ScienceQA—55.8%

Multilingual Llama 2-13B leads

Gemma 1.1 2b IT: 24.6 (#289), Llama 2-13B: 26.5 (#279)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Non-English9881024
LMArena Chinese10121001
LMArena German9441009
LMArena Korean899953
LMArena Russian9901055
LMArena French—1044
LMArena Japanese—894
LMArena Spanish—1087

Instruction Following Llama 2-13B leads

Gemma 1.1 2b IT: 49.9 (#299), Llama 2-13B: 53.3 (#287)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Instruction Following9921045

Long Context Llama 2-13B leads

Gemma 1.1 2b IT: 30.6 (#286), Llama 2-13B: 32.3 (#269)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Longer Query10031064

Writing & Preference Llama 2-13B leads

Gemma 1.1 2b IT: 25.1 (#306), Llama 2-13B: 29.8 (#289)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITLlama 2-13B
LMArena Text10221084
LMArena Creative Writing9981047
LMArena Multi-Turn9591050

Frequently asked questions

Is Gemma 1.1 2b IT better than Llama 2-13B?

Gemma 1.1 2b IT and Llama 2-13B score almost the same on the Noometry Index (29.3 vs 29.6), so choose on price, context window or the category you care about most.

Is Gemma 1.1 2b IT or Llama 2-13B better for coding?

They score almost the same on coding (30.1 vs 30.9); test both on your own repository before choosing.

How many benchmarks do Gemma 1.1 2b IT and Llama 2-13B share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Llama 2-13B has 32.

Related comparisons

Go deeper