Model comparison

Codellama 34b Instruct vs Llama 3.1-8B

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 23.0 on the Noometry Index.

Last verified . 14 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Llama 3.1-8B Meta

23.0

Rank #352 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Codellama 34b Instruct scores higher in 3 categories and Llama 3.1-8B in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Codellama 34b Instruct leads 31.0 to 10.2.

Side by side

Codellama 34b Instruct and Llama 3.1-8B specifications
Codellama 34b InstructLlama 3.1-8B
ProviderMetaMeta
Noometry Index30.823.0
Released—2024-07-23
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.08
Results tracked1443

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 34b Instruct leads

Codellama 34b Instruct: 28.5 (#314), Llama 3.1-8B: 20.2 (#340)

Coding benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
BigCodeBench Instruct29%32.8%
LMArena Coding10461195
BigCodeBench Complete37.1%40.5%
HumanEval+43.9%62.8%
MBPP+56.3%55.6%
SciCode—13.2%
WeirdML—1.7%

Agentic & Tool Use Not comparable

Codellama 34b Instruct: —, Llama 3.1-8B: 22.5 (#131)

Agentic & Tool Use benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
Berkeley Function Calling Leaderboard—25.8%
BALROG—15.1%

Reasoning Codellama 34b Instruct leads

Codellama 34b Instruct: 19.6 (#255), Llama 3.1-8B: 14.9 (#321)

Reasoning benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
LMArena Hard Prompts10321175
CritPt—0%
Chess Puzzles—0%
DTBench—50.9%
LMCA—5.4%
Epoch Capabilities Index—116.57
PIQA—81.2%

Math Codellama 34b Instruct leads

Codellama 34b Instruct: 31.0 (#230), Llama 3.1-8B: 10.2 (#317)

Math benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
LMArena Math10561179
OTIS Mock AIME 2024-2025—1.7%
Omni-MATH—13.7%
MATH Level 5—22.9%
GSM8K—82.4%

Knowledge Not comparable

Codellama 34b Instruct: —, Llama 3.1-8B: 8.0 (#307)

Knowledge benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
GPQA Diamond—27%
MMLU-Pro—40.6%
GPQA (HELM)—24.7%
LMArena Expert—1144
BoolQ—82.8%
MMLU—56.1%

Multilingual Llama 3.1-8B leads

Codellama 34b Instruct: 25.8 (#284), Llama 3.1-8B: 34.0 (#249)

Multilingual benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
LMArena Non-English10111148
LMArena Chinese9761151
LMArena French—1177
LMArena German—1144
LMArena Japanese—1061
LMArena Korean—1053
LMArena Russian—1158
LMArena Spanish—1169

Instruction Following Llama 3.1-8B leads

Codellama 34b Instruct: 52.2 (#291), Llama 3.1-8B: 58.9 (#258)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
LMArena Instruction Following10281159
IFEval—74.3%

Long Context Llama 3.1-8B leads

Codellama 34b Instruct: 30.9 (#284), Llama 3.1-8B: 35.8 (#238)

Long Context benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
LMArena Longer Query10131182

Writing & Preference Llama 3.1-8B leads

Codellama 34b Instruct: 28.2 (#297), Llama 3.1-8B: 29.7 (#290)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructLlama 3.1-8B
LMArena Text10661187
LMArena Creative Writing10321154
LMArena Multi-Turn10151172
EQ-Bench Creative Writing—713
WildBench—68.7%

Frequently asked questions

Is Codellama 34b Instruct better than Llama 3.1-8B?

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 23.0 on the Noometry Index.

Is Codellama 34b Instruct or Llama 3.1-8B better for coding?

Codellama 34b Instruct scores higher on coding benchmarks: 28.5 versus 20.2 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Llama 3.1-8B share?

14 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Llama 3.1-8B has 43.

Related comparisons

Go deeper