Model comparison

Codellama 70b Instruct vs Magistral Small

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 30.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • The widest gap is in reasoning, where Codellama 70b Instruct leads 20.1 to 6.8.

Side by side

Codellama 70b Instruct and Magistral Small specifications
Codellama 70b InstructMagistral Small
ProviderMetaMistral AI
Noometry Index33.730.2
Released—2025-06-10
WeightsOpenOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Codellama 70b Instruct: 37.6 (#193), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkCodellama 70b InstructMagistral Small
SciCode—35.2%
BigCodeBench Instruct40.7%—
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Reasoning Codellama 70b Instruct leads

Codellama 70b Instruct: 20.1 (#242), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkCodellama 70b InstructMagistral Small
ARC-AGI-2—0%
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
Chess Puzzles—3%
LMArena Hard Prompts1052—
DTBench—61.3%
Epoch Capabilities Index—133.19

Math Not comparable

Codellama 70b Instruct: —, Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkCodellama 70b InstructMagistral Small
OTIS Mock AIME 2024-2025—30%

Knowledge Not comparable

Codellama 70b Instruct: —, Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkCodellama 70b InstructMagistral Small
GPQA Diamond—56.1%

Multilingual Not comparable

Codellama 70b Instruct: 24.8 (#288), Magistral Small: —

Multilingual benchmarks
BenchmarkCodellama 70b InstructMagistral Small
LMArena Non-English992—

Instruction Following Not comparable

Codellama 70b Instruct: 51.9 (#293), Magistral Small: —

Instruction Following benchmarks
BenchmarkCodellama 70b InstructMagistral Small
LMArena Instruction Following1024—

Writing & Preference Not comparable

Codellama 70b Instruct: 33.4 (#277), Magistral Small: —

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructMagistral Small
LMArena Text1057—

Frequently asked questions

Is Codellama 70b Instruct better than Magistral Small?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 30.2 on the Noometry Index.

Is Codellama 70b Instruct or Magistral Small better for coding?

They score almost the same on coding (37.6 vs 38.4); test both on your own repository before choosing.

How many benchmarks do Codellama 70b Instruct and Magistral Small share?

0 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper