Model comparison

Codellama 70b Instruct vs Granite 4.1 8b

Granite 4.1 8b is the stronger model overall, scoring 37.4 to 33.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Granite 4.1 8b IBM

37.4

Rank #205 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 1 category and Granite 4.1 8b in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Granite 4.1 8b leads 41.7 to 24.8.

Side by side

Codellama 70b Instruct and Granite 4.1 8b specifications
Codellama 70b InstructGranite 4.1 8b
ProviderMetaIBM
Noometry Index33.737.4
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked713

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Granite 4.1 8b: 30.1 (#297)

Coding benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena WebDev—1192
BigCodeBench Instruct40.7%—
LMArena Coding—1312
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Reasoning Granite 4.1 8b leads

Codellama 70b Instruct: 20.1 (#242), Granite 4.1 8b: 25.7 (#143)

Reasoning benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena Hard Prompts10521293

Math Not comparable

Codellama 70b Instruct: —, Granite 4.1 8b: 36.4 (#166)

Math benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena Math—1312

Knowledge Not comparable

Codellama 70b Instruct: —, Granite 4.1 8b: 36.1 (#174)

Knowledge benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena Expert—1309

Multilingual Granite 4.1 8b leads

Codellama 70b Instruct: 24.8 (#288), Granite 4.1 8b: 41.7 (#204)

Multilingual benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena Non-English9921261
LMArena Chinese—1337
LMArena Russian—1240

Instruction Following Granite 4.1 8b leads

Codellama 70b Instruct: 51.9 (#293), Granite 4.1 8b: 66.9 (#203)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena Instruction Following10241269

Long Context Not comparable

Codellama 70b Instruct: —, Granite 4.1 8b: 38.7 (#193)

Long Context benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena Longer Query—1275

Writing & Preference Granite 4.1 8b leads

Codellama 70b Instruct: 33.4 (#277), Granite 4.1 8b: 48.1 (#204)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructGranite 4.1 8b
LMArena Text10571290
LMArena Creative Writing—1251
LMArena Multi-Turn—1268

Frequently asked questions

Is Codellama 70b Instruct better than Granite 4.1 8b?

Granite 4.1 8b is the stronger model overall, scoring 37.4 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or Granite 4.1 8b better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 30.1 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Granite 4.1 8b share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Granite 4.1 8b has 13.

Related comparisons

Go deeper