Model comparison

Granite 3.0 8b Instruct vs Llama2 70b Steerlm Chat

Granite 3.0 8b Instruct and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (31.6 vs 31.8), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

Granite 3.0 8b Instruct IBM

31.6

Rank #270 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Granite 3.0 8b Instruct scores higher in 4 categories and Llama2 70b Steerlm Chat in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Granite 3.0 8b Instruct leads 34.0 to 30.4.

Side by side

Granite 3.0 8b Instruct and Llama2 70b Steerlm Chat specifications
Granite 3.0 8b InstructLlama2 70b Steerlm Chat
ProviderIBMNVIDIA
Noometry Index31.631.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked149

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Granite 3.0 8b Instruct: 29.7 (#301), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Coding11121025
BigCodeBench Instruct29.3%—
BigCodeBench Complete35.4%—

Reasoning Too close to call

Granite 3.0 8b Instruct: 20.9 (#230), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Hard Prompts10921047

Math Granite 3.0 8b Instruct leads

Granite 3.0 8b Instruct: 32.8 (#210), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Math11431072

Knowledge Not comparable

Granite 3.0 8b Instruct: 29.6 (#236), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Expert1087—

Multilingual Llama2 70b Steerlm Chat leads

Granite 3.0 8b Instruct: 27.2 (#276), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Non-English10371063
LMArena Chinese1063—
LMArena Russian1060—

Instruction Following Granite 3.0 8b Instruct leads

Granite 3.0 8b Instruct: 56.0 (#276), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Instruction Following10881060

Long Context Granite 3.0 8b Instruct leads

Granite 3.0 8b Instruct: 34.0 (#252), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Longer Query1122998

Writing & Preference Too close to call

Granite 3.0 8b Instruct: 31.1 (#285), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkGranite 3.0 8b InstructLlama2 70b Steerlm Chat
LMArena Text10961098
LMArena Creative Writing10711091
LMArena Multi-Turn10631058

Frequently asked questions

Is Granite 3.0 8b Instruct better than Llama2 70b Steerlm Chat?

Granite 3.0 8b Instruct and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (31.6 vs 31.8), so choose on price, context window or the category you care about most.

Is Granite 3.0 8b Instruct or Llama2 70b Steerlm Chat better for coding?

They score almost the same on coding (29.7 vs 29.9); test both on your own repository before choosing.

How many benchmarks do Granite 3.0 8b Instruct and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. Granite 3.0 8b Instruct has 14 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper