Model comparison

Codellama 70b Instruct vs Gemma 7B

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 30.0 on the Noometry Index.

Last verified . 5 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Gemma 7B Google

30.0

Rank #299 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Codellama 70b Instruct scores higher in 4 categories and Gemma 7B in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Codellama 70b Instruct leads 37.6 to 30.5.

Side by side

Codellama 70b Instruct and Gemma 7B specifications
Codellama 70b InstructGemma 7B
ProviderMetaGoogle
Noometry Index33.730.0
Released—2024-02-21
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked727

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Gemma 7B: 30.5 (#294)

Coding benchmarks
BenchmarkCodellama 70b InstructGemma 7B
HumanEval+65.9%28.7%
BigCodeBench Instruct40.7%—
LMArena Coding—1048
BigCodeBench Complete49.6%—
MBPP+—43.4%

Reasoning Too close to call

Codellama 70b Instruct: 20.1 (#242), Gemma 7B: 19.9 (#249)

Reasoning benchmarks
BenchmarkCodellama 70b InstructGemma 7B
LMArena Hard Prompts10521042
Adversarial NLI—48.7%
BIG-Bench Hard—55.1%
Epoch Capabilities Index—111.99
HellaSwag—82.2%
PIQA—81.2%
WinoGrande—79%

Math Not comparable

Codellama 70b Instruct: —, Gemma 7B: 31.2 (#228)

Math benchmarks
BenchmarkCodellama 70b InstructGemma 7B
LMArena Math—1066
GSM8K—46.4%

Knowledge Not comparable

Codellama 70b Instruct: —, Gemma 7B: 27.3 (#252)

Knowledge benchmarks
BenchmarkCodellama 70b InstructGemma 7B
LMArena Expert—1001
ARC (AI2) Challenge—78.3%
BoolQ—83.2%
MMLU—66.1%
OpenBookQA—78.6%
TriviaQA—72.3%

Multilingual Too close to call

Codellama 70b Instruct: 24.8 (#288), Gemma 7B: 25.1 (#287)

Multilingual benchmarks
BenchmarkCodellama 70b InstructGemma 7B
LMArena Non-English992999
LMArena Chinese—1035
LMArena French—1025
LMArena Russian—993

Instruction Following Too close to call

Codellama 70b Instruct: 51.9 (#293), Gemma 7B: 51.5 (#295)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructGemma 7B
LMArena Instruction Following10241017

Long Context Not comparable

Codellama 70b Instruct: —, Gemma 7B: 31.1 (#282)

Long Context benchmarks
BenchmarkCodellama 70b InstructGemma 7B
LMArena Longer Query—1022

Writing & Preference Codellama 70b Instruct leads

Codellama 70b Instruct: 33.4 (#277), Gemma 7B: 27.1 (#302)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructGemma 7B
LMArena Text10571056
LMArena Creative Writing—1024
LMArena Multi-Turn—963

Frequently asked questions

Is Codellama 70b Instruct better than Gemma 7B?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 30.0 on the Noometry Index.

Is Codellama 70b Instruct or Gemma 7B better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 30.5 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Gemma 7B share?

5 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Gemma 7B has 27.

Related comparisons

Go deeper