Model comparison

Codellama 34b Instruct vs Mixtral 8x22B

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 27.1 on the Noometry Index.

Last verified . 14 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Codellama 34b Instruct scores higher in 2 categories and Mixtral 8x22B in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mixtral 8x22B leads 36.9 to 28.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 37.1% for Codellama 34b Instruct and 50.2% for Mixtral 8x22B.

Side by side

Codellama 34b Instruct and Mixtral 8x22B specifications
Codellama 34b InstructMixtral 8x22B
ProviderMetaMistral AI
Noometry Index30.827.1
Released—2024-04-17
WeightsOpenOpen
Context window—64K
Max output—64K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1434

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 34b Instruct leads

Codellama 34b Instruct: 28.5 (#314), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
BigCodeBench Instruct29%40.6%
LMArena Coding10461166
BigCodeBench Complete37.1%50.2%
HumanEval+43.9%72%
MBPP+56.3%64.3%
WeirdML—3.2%

Agentic & Tool Use Not comparable

Codellama 34b Instruct: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
Cybench—7.5%

Reasoning Too close to call

Codellama 34b Instruct: 19.6 (#255), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
LMArena Hard Prompts10321150
DTBench—55.1%
Epoch Capabilities Index—122.03
ForecastBench—56.3

Math Codellama 34b Instruct leads

Codellama 34b Instruct: 31.0 (#230), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
LMArena Math10561184
Omni-MATH—16.3%
MATH Level 5—24.2%

Knowledge Not comparable

Codellama 34b Instruct: —, Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
MMLU—77.8%

Multilingual Mixtral 8x22B leads

Codellama 34b Instruct: 25.8 (#284), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
LMArena Non-English10111128
LMArena Chinese9761116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Mixtral 8x22B leads

Codellama 34b Instruct: 52.2 (#291), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
LMArena Instruction Following10281147
IFEval—72.4%

Long Context Mixtral 8x22B leads

Codellama 34b Instruct: 30.9 (#284), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
LMArena Longer Query10131144

Writing & Preference Mixtral 8x22B leads

Codellama 34b Instruct: 28.2 (#297), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructMixtral 8x22B
LMArena Text10661162
LMArena Creative Writing10321141
LMArena Multi-Turn10151130
WildBench—71.1%

Frequently asked questions

Is Codellama 34b Instruct better than Mixtral 8x22B?

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 27.1 on the Noometry Index.

Is Codellama 34b Instruct or Mixtral 8x22B better for coding?

Codellama 34b Instruct scores higher on coding benchmarks: 28.5 versus 24.2 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Mixtral 8x22B share?

14 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper