Model comparison

Falcon-180B vs Granite 3.1 8b Instruct

Falcon-180B and Granite 3.1 8b Instruct score almost the same on the Noometry Index (32.2 vs 32.4), so choose on price, context window or the category you care about most.

Last verified . 6 shared benchmarks.

Granite 3.1 8b Instruct IBM

32.4

Rank #258 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Granite 3.1 8b Instruct in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Granite 3.1 8b Instruct leads 35.5 to 29.1.

Side by side

Falcon-180B and Granite 3.1 8b Instruct specifications
Falcon-180BGranite 3.1 8b Instruct
ProviderTechnology Innovation InstituteIBM
Noometry Index32.232.4
Released2023-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1613

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Granite 3.1 8b Instruct: 34.5 (#233)

Coding benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Coding—1186

Agentic & Tool Use Not comparable

Falcon-180B: —, Granite 3.1 8b Instruct: 24.1 (#120)

Agentic & Tool Use benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
Berkeley Function Calling Leaderboard—27.1%

Reasoning Granite 3.1 8b Instruct leads

Falcon-180B: 19.1 (#269), Granite 3.1 8b Instruct: 22.1 (#207)

Reasoning benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Hard Prompts10071145
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Granite 3.1 8b Instruct: 33.0 (#209)

Math benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Math—1152
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Granite 3.1 8b Instruct: 31.1 (#220)

Knowledge benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Expert—1142
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Granite 3.1 8b Instruct leads

Falcon-180B: 25.2 (#286), Granite 3.1 8b Instruct: 30.9 (#260)

Multilingual benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Non-English10001099
LMArena Chinese—1145
LMArena Russian—1092

Instruction Following Granite 3.1 8b Instruct leads

Falcon-180B: 53.4 (#286), Granite 3.1 8b Instruct: 58.6 (#259)

Instruction Following benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Instruction Following10471131

Long Context Not comparable

Falcon-180B: —, Granite 3.1 8b Instruct: 35.2 (#241)

Long Context benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Longer Query—1162

Writing & Preference Granite 3.1 8b Instruct leads

Falcon-180B: 29.1 (#295), Granite 3.1 8b Instruct: 35.5 (#266)

Writing & Preference benchmarks
BenchmarkFalcon-180BGranite 3.1 8b Instruct
LMArena Text10541150
LMArena Creative Writing10891129
LMArena Multi-Turn10131108

Frequently asked questions

Is Falcon-180B better than Granite 3.1 8b Instruct?

Falcon-180B and Granite 3.1 8b Instruct score almost the same on the Noometry Index (32.2 vs 32.4), so choose on price, context window or the category you care about most.

How many benchmarks do Falcon-180B and Granite 3.1 8b Instruct share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Granite 3.1 8b Instruct has 13.

Related comparisons

Go deeper