Model comparison

Granite 3.1 8b Instruct vs Olmo 2 0325 32b Instruct

Granite 3.1 8b Instruct and Olmo 2 0325 32b Instruct score almost the same on the Noometry Index (32.4 vs 32.7), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Summary

  • They share 11 benchmarks with published results for both. Granite 3.1 8b Instruct scores higher in 2 categories and Olmo 2 0325 32b Instruct in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Granite 3.1 8b Instruct leads 31.1 to 19.5.

Side by side

Granite 3.1 8b Instruct and Olmo 2 0325 32b Instruct specifications
Granite 3.1 8b InstructOlmo 2 0325 32b Instruct
ProviderIBMAllen Institute for AI (Ai2)
Noometry Index32.432.7
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1316

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Granite 3.1 8b Instruct: 34.5 (#233), Olmo 2 0325 32b Instruct: 35.2 (#227)

Coding benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
LMArena Coding11861210

Agentic & Tool Use Not comparable

Granite 3.1 8b Instruct: 24.1 (#120), Olmo 2 0325 32b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
Berkeley Function Calling Leaderboard27.1%—

Reasoning Olmo 2 0325 32b Instruct leads

Granite 3.1 8b Instruct: 22.1 (#207), Olmo 2 0325 32b Instruct: 23.6 (#175)

Reasoning benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
LMArena Hard Prompts11451208

Math Granite 3.1 8b Instruct leads

Granite 3.1 8b Instruct: 33.0 (#209), Olmo 2 0325 32b Instruct: 26.8 (#255)

Math benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
LMArena Math11521208
Omni-MATH—16.1%

Knowledge Granite 3.1 8b Instruct leads

Granite 3.1 8b Instruct: 31.1 (#220), Olmo 2 0325 32b Instruct: 19.5 (#279)

Knowledge benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
MMLU-Pro—41.4%
GPQA (HELM)—28.7%
LMArena Expert1142—

Multilingual Olmo 2 0325 32b Instruct leads

Granite 3.1 8b Instruct: 30.9 (#260), Olmo 2 0325 32b Instruct: 34.8 (#248)

Multilingual benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
LMArena Non-English10991160
LMArena Chinese11451192
LMArena Russian10921187

Instruction Following Olmo 2 0325 32b Instruct leads

Granite 3.1 8b Instruct: 58.6 (#259), Olmo 2 0325 32b Instruct: 61.5 (#244)

Instruction Following benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
LMArena Instruction Following11311186
IFEval—78%

Long Context Too close to call

Granite 3.1 8b Instruct: 35.2 (#241), Olmo 2 0325 32b Instruct: 36.2 (#234)

Long Context benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
LMArena Longer Query11621194

Writing & Preference Olmo 2 0325 32b Instruct leads

Granite 3.1 8b Instruct: 35.5 (#266), Olmo 2 0325 32b Instruct: 42.1 (#236)

Writing & Preference benchmarks
BenchmarkGranite 3.1 8b InstructOlmo 2 0325 32b Instruct
LMArena Text11501218
LMArena Creative Writing11291199
LMArena Multi-Turn11081221
WildBench—73.4%

Frequently asked questions

Is Granite 3.1 8b Instruct better than Olmo 2 0325 32b Instruct?

Granite 3.1 8b Instruct and Olmo 2 0325 32b Instruct score almost the same on the Noometry Index (32.4 vs 32.7), so choose on price, context window or the category you care about most.

Is Granite 3.1 8b Instruct or Olmo 2 0325 32b Instruct better for coding?

They score almost the same on coding (34.5 vs 35.2); test both on your own repository before choosing.

How many benchmarks do Granite 3.1 8b Instruct and Olmo 2 0325 32b Instruct share?

11 benchmarks have published results for both models. Granite 3.1 8b Instruct has 13 scored results on Noometry and Olmo 2 0325 32b Instruct has 16.

Related comparisons

Go deeper