Model comparison

Claude 3 Haiku vs Mistral Nemo

Claude 3 Haiku and Mistral Nemo score almost the same on the Noometry Index (25.9 vs 26.4), so choose on price, context window or the category you care about most.

Last verified . 5 shared benchmarks.

Claude 3 Haiku Anthropic

25.9

Rank #340 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Claude 3 Haiku scores higher in 2 categories and Mistral Nemo in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Nemo leads 25.5 to 9.8.
  • The biggest single-benchmark swing is GPQA Diamond: 36.3% for Claude 3 Haiku and 29.9% for Mistral Nemo.
  • Mistral Nemo has downloadable open weights; the other is API-only.

Side by side

Claude 3 Haiku and Mistral Nemo specifications
Claude 3 HaikuMistral Nemo
ProviderAnthropicMistral AI
Noometry Index25.926.4
Released2024-03-072024-07-01
WeightsProprietaryOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked3710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 3 Haiku: 26.4 (#325), Mistral Nemo: —

Coding benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
WeirdML9.8%—
BigCodeBench Instruct39.4%—
LMArena Coding1199—
BigCodeBench Complete50.1%—
CadEval12%—
HumanEval+68.9%—
MBPP+68.8%—

Agentic & Tool Use Not comparable

Claude 3 Haiku: —, Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%

Reasoning Mistral Nemo leads

Claude 3 Haiku: 16.3 (#307), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
DTBench50.1%48.6%
Epoch Capabilities Index118.35118.68
Kagi LLM Benchmark34.2%—
LMArena Hard Prompts1174—
LMCA8.8%—
ForecastBench53.2—
PIQA—83.5%
WinoGrande74.2%—

Math Mistral Nemo leads

Claude 3 Haiku: 9.8 (#319), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
MATH Level 514.9%10.8%
OTIS Mock AIME 2024-20251.8%—
LMArena Math1188—
GSM8K—84.2%

Knowledge Claude 3 Haiku leads

Claude 3 Haiku: 17.3 (#285), Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
GPQA Diamond36.3%29.9%
Confabulations34.2%—
LMArena Expert1148—
BoolQ—82.5%
MMLU73.8%—

Multimodal Not comparable

Claude 3 Haiku: 23.6 (#128), Mistral Nemo: —

Multimodal benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
LMArena Vision950—
ScienceQA72%—

Multilingual Not comparable

Claude 3 Haiku: 36.0 (#243), Mistral Nemo: —

Multilingual benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
LMArena Non-English1178—
LMArena Chinese1155—
LMArena French1195—
LMArena German1174—
LMArena Japanese1102—
LMArena Korean1109—
LMArena Russian1204—
LMArena Spanish1166—

Instruction Following Not comparable

Claude 3 Haiku: 61.3 (#247), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
LMArena Instruction Following1173—

Long Context Not comparable

Claude 3 Haiku: 36.1 (#237), Mistral Nemo: —

Long Context benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
LMArena Longer Query1190—

Writing & Preference Claude 3 Haiku leads

Claude 3 Haiku: 29.7 (#291), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkClaude 3 HaikuMistral Nemo
EQ-Bench Creative Writing717881
LMArena Text1195—
LMArena Creative Writing1157—
LMArena Multi-Turn1190—

Frequently asked questions

Is Claude 3 Haiku better than Mistral Nemo?

Claude 3 Haiku and Mistral Nemo score almost the same on the Noometry Index (25.9 vs 26.4), so choose on price, context window or the category you care about most.

How many benchmarks do Claude 3 Haiku and Mistral Nemo share?

5 benchmarks have published results for both models. Claude 3 Haiku has 37 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper