Model comparison

Granite 3.1 2b Instruct vs Mercury 2.5

Granite 3.1 2b Instruct and Mercury 2.5 score almost the same on the Noometry Index (33.2 vs 33.5), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Granite 3.1 2b Instruct IBM

33.2

Rank #247 Confirmed

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • The widest gap is in math, where Granite 3.1 2b Instruct leads 33.1 to 23.3.
  • Granite 3.1 2b Instruct has downloadable open weights; the other is API-only.

Side by side

Granite 3.1 2b Instruct and Mercury 2.5 specifications
Granite 3.1 2b InstructMercury 2.5
ProviderIBMInception
Noometry Index33.233.5
Released—2026-09-08
WeightsOpenProprietary
Context window—260K
Max output—66K
Input $ / M tokens—$0.04
Output $ / M tokens—$0.15
Results tracked124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Granite 3.1 2b Instruct: 33.4 (#257), Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
SciCode—38.5%
LMArena Coding1149—
ALE-Bench—301.65

Reasoning Too close to call

Granite 3.1 2b Instruct: 22.0 (#209), Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
CritPt—0%
LMArena Hard Prompts1138—

Math Granite 3.1 2b Instruct leads

Granite 3.1 2b Instruct: 33.1 (#206), Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
ProofBench—3%
LMArena Math1159—

Knowledge Not comparable

Granite 3.1 2b Instruct: 30.8 (#224), Mercury 2.5: —

Knowledge benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
LMArena Expert1131—

Multilingual Not comparable

Granite 3.1 2b Instruct: 29.1 (#269), Mercury 2.5: —

Multilingual benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
LMArena Non-English1068—
LMArena Chinese1139—
LMArena Russian1063—

Instruction Following Not comparable

Granite 3.1 2b Instruct: 57.7 (#264), Mercury 2.5: —

Instruction Following benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
LMArena Instruction Following1116—

Long Context Not comparable

Granite 3.1 2b Instruct: 35.0 (#244), Mercury 2.5: —

Long Context benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
LMArena Longer Query1155—

Writing & Preference Not comparable

Granite 3.1 2b Instruct: 34.1 (#274), Mercury 2.5: —

Writing & Preference benchmarks
BenchmarkGranite 3.1 2b InstructMercury 2.5
LMArena Text1127—
LMArena Creative Writing1116—
LMArena Multi-Turn1099—

Frequently asked questions

Is Granite 3.1 2b Instruct better than Mercury 2.5?

Granite 3.1 2b Instruct and Mercury 2.5 score almost the same on the Noometry Index (33.2 vs 33.5), so choose on price, context window or the category you care about most.

Is Granite 3.1 2b Instruct or Mercury 2.5 better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 33.4 in the Noometry coding category.

How many benchmarks do Granite 3.1 2b Instruct and Mercury 2.5 share?

0 benchmarks have published results for both models. Granite 3.1 2b Instruct has 12 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper