Model comparison

Claude 2.1 vs Falcon-180B

Falcon-180B is the stronger model overall, scoring 32.2 to 25.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 1 category and Falcon-180B in 0 categories; one gap is clear of the uncertainty.
  • Falcon-180B has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Falcon-180B specifications
Claude 2.1Falcon-180B
ProviderAnthropicTechnology Innovation Institute
Noometry Index25.232.2
Released2023-11-212023-09-06
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked716

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2.1: 26.2 (#327), Falcon-180B: —

Coding benchmarks
BenchmarkClaude 2.1Falcon-180B
WeirdML7.1%—

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), Falcon-180B: 19.1 (#269)

Reasoning benchmarks
BenchmarkClaude 2.1Falcon-180B
Epoch Capabilities Index119.27112.13
LMArena Hard Prompts—1007
DTBench51%—
ForecastBench54.2—
HellaSwag—89%
LAMBADA—79.8%
PIQA—84.9%
WinoGrande—87.1%

Math Not comparable

Claude 2.1: 10.2 (#315), Falcon-180B: —

Math benchmarks
BenchmarkClaude 2.1Falcon-180B
OTIS Mock AIME 2024-20251.9%—
GSM8K—54.4%

Knowledge Not comparable

Claude 2.1: 15.4 (#292), Falcon-180B: —

Knowledge benchmarks
BenchmarkClaude 2.1Falcon-180B
MMLU73.5%70.6%
GPQA Diamond33%—
ARC (AI2) Challenge—67.8%
BoolQ—89%
OpenBookQA—64.2%

Multilingual Not comparable

Claude 2.1: —, Falcon-180B: 25.2 (#286)

Multilingual benchmarks
BenchmarkClaude 2.1Falcon-180B
LMArena Non-English—1000

Instruction Following Not comparable

Claude 2.1: —, Falcon-180B: 53.4 (#286)

Instruction Following benchmarks
BenchmarkClaude 2.1Falcon-180B
LMArena Instruction Following—1047

Writing & Preference Not comparable

Claude 2.1: —, Falcon-180B: 29.1 (#295)

Writing & Preference benchmarks
BenchmarkClaude 2.1Falcon-180B
LMArena Text—1054
LMArena Creative Writing—1089
LMArena Multi-Turn—1013

Frequently asked questions

Is Claude 2.1 better than Falcon-180B?

Falcon-180B is the stronger model overall, scoring 32.2 to 25.2 on the Noometry Index.

How many benchmarks do Claude 2.1 and Falcon-180B share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Falcon-180B has 16.

Related comparisons

Go deeper