Model comparison

Claude 2.1 vs Mistral

Mistral is the stronger model overall, scoring 29.9 to 25.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • The widest gap is in math, where Mistral leads 22.3 to 10.2.

Side by side

Claude 2.1 and Mistral specifications
Claude 2.1Mistral
ProviderAnthropicMistral AI
Noometry Index25.229.9
Released2023-11-21—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked722

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Claude 2.1: 26.2 (#327), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkClaude 2.1Mistral
WeirdML7.1%—
LMArena Coding—1162

Reasoning Too close to call

Claude 2.1: 21.4 (#221), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkClaude 2.1Mistral
LMArena Hard Prompts—1149
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math Mistral leads

Claude 2.1: 10.2 (#315), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkClaude 2.1Mistral
OTIS Mock AIME 2024-20251.9%—
Omni-MATH—7.2%
LMArena Math—1180

Knowledge Mistral leads

Claude 2.1: 15.4 (#292), Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkClaude 2.1Mistral
GPQA Diamond33%—
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
LMArena Expert—1125
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkClaude 2.1Mistral
LMArena Non-English—1129
LMArena Chinese—1109
LMArena French—1180
LMArena German—1155
LMArena Japanese—1013
LMArena Korean—1032
LMArena Russian—1168
LMArena Spanish—1143

Instruction Following Not comparable

Claude 2.1: —, Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkClaude 2.1Mistral
IFEval—56.8%
LMArena Instruction Following—1152

Long Context Not comparable

Claude 2.1: —, Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkClaude 2.1Mistral
LMArena Longer Query—1153

Writing & Preference Not comparable

Claude 2.1: —, Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkClaude 2.1Mistral
LMArena Text—1165
LMArena Creative Writing—1158
WildBench—66%
LMArena Multi-Turn—1147

Frequently asked questions

Is Claude 2.1 better than Mistral?

Mistral is the stronger model overall, scoring 29.9 to 25.2 on the Noometry Index.

Is Claude 2.1 or Mistral better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Mistral share?

0 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper