Model comparison

Falcon-180B vs Mistral Large

Falcon-180B and Mistral Large score almost the same on the Noometry Index (32.2 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 8 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Falcon-180B scores higher in 1 category and Mistral Large in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral Large leads 40.0 to 25.2.

Side by side

Falcon-180B and Mistral Large specifications
Falcon-180BMistral Large
ProviderTechnology Innovation InstituteMistral AI
Noometry Index32.231.9
Released2023-09-062024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1651

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkFalcon-180BMistral Large
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Falcon-180B: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkFalcon-180BMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Falcon-180B leads

Falcon-180B: 19.1 (#269), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkFalcon-180BMistral Large
LMArena Hard Prompts10071257
Epoch Capabilities Index112.13128.52
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
ForecastBench—57.1
HellaSwag89%—
LAMBADA79.8%—
LiveBench—48.4%
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkFalcon-180BMistral Large
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkFalcon-180BMistral Large
MMLU70.6%80%
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
ARC (AI2) Challenge67.8%—
BoolQ89%—
OpenBookQA64.2%—

Multilingual Mistral Large leads

Falcon-180B: 25.2 (#286), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkFalcon-180BMistral Large
LMArena Non-English10001237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Mistral Large leads

Falcon-180B: 53.4 (#286), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkFalcon-180BMistral Large
LMArena Instruction Following10471249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Not comparable

Falcon-180B: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkFalcon-180BMistral Large
LMArena Longer Query—1261

Writing & Preference Mistral Large leads

Falcon-180B: 29.1 (#295), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkFalcon-180BMistral Large
LMArena Text10541266
LMArena Creative Writing10891243
LMArena Multi-Turn10131260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Falcon-180B better than Mistral Large?

Falcon-180B and Mistral Large score almost the same on the Noometry Index (32.2 vs 31.9), so choose on price, context window or the category you care about most.

How many benchmarks do Falcon-180B and Mistral Large share?

8 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper