Model comparison

Claude 3.5 Haiku vs Phi 3 Small 8k Instruct

Claude 3.5 Haiku and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.2 vs 29.3), so choose on price, context window or the category you care about most.

Last verified . 25 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 6 categories and Phi 3 Small 8k Instruct in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Phi 3 Small 8k Instruct leads 27.6 to 14.7.
  • The biggest single-benchmark swing is LiveBench Coding: 51.4% for Claude 3.5 Haiku and 20.3% for Phi 3 Small 8k Instruct.
  • Phi 3 Small 8k Instruct has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Phi 3 Small 8k Instruct specifications
Claude 3.5 HaikuPhi 3 Small 8k Instruct
ProviderAnthropicMicrosoft
Noometry Index29.229.3
Released2024-10-222024-04-23
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4932

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Haiku leads

Claude 3.5 Haiku: 32.9 (#265), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LiveBench Coding51.4%20.3%
LMArena Coding12861101
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Phi 3 Small 8k Instruct: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
BALROG19.3%—

Reasoning Claude 3.5 Haiku leads

Claude 3.5 Haiku: 17.7 (#290), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LiveBench Reasoning28.1%15.9%
LMArena Hard Prompts12511100
LiveBench Data Analysis48.5%30.3%
LiveBench43.5%24%
CritPt0%—
DTBench56.7%—
Adversarial NLI—58.1%
BIG-Bench Hard—79.1%
Epoch Capabilities Index127.15—
HellaSwag—77%
WinoGrande—81.5%

Math Phi 3 Small 8k Instruct leads

Claude 3.5 Haiku: 14.7 (#300), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LiveBench Math35.5%17.6%
LMArena Math12441151
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Phi 3 Small 8k Instruct leads

Claude 3.5 Haiku: 18.7 (#281), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LMArena Expert12081067
MMLU74.3%75.7%
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
ARC (AI2) Challenge—90.7%
OpenBookQA—88%
TriviaQA—58.1%

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Phi 3 Small 8k Instruct: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LMArena Vision1092—
GeoBench34%—

Multilingual Claude 3.5 Haiku leads

Claude 3.5 Haiku: 40.0 (#218), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LMArena Non-English12381058
LMArena Chinese12291061
LMArena French12641135
LMArena German12371080
LMArena Japanese1175966
LMArena Korean1173894
LMArena Russian12531111
LMArena Spanish12611111

Instruction Following Claude 3.5 Haiku leads

Claude 3.5 Haiku: 62.9 (#234), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LiveBench Instruction Following61.9%47.2%
LMArena Instruction Following12411087
IFEval79.2%—

Long Context Claude 3.5 Haiku leads

Claude 3.5 Haiku: 38.3 (#200), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LMArena Longer Query12611088

Writing & Preference Claude 3.5 Haiku leads

Claude 3.5 Haiku: 42.7 (#234), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuPhi 3 Small 8k Instruct
LMArena Text12551110
LMArena Creative Writing12331083
LMArena Multi-Turn12651068
LiveBench Language35.4%12.9%
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—

Frequently asked questions

Is Claude 3.5 Haiku better than Phi 3 Small 8k Instruct?

Claude 3.5 Haiku and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.2 vs 29.3), so choose on price, context window or the category you care about most.

Is Claude 3.5 Haiku or Phi 3 Small 8k Instruct better for coding?

Claude 3.5 Haiku scores higher on coding benchmarks: 32.9 versus 27.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Phi 3 Small 8k Instruct share?

25 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper