Model comparison

Claude 3.5 Haiku vs GPT-4o

Claude 3.5 Haiku and GPT-4o score almost the same on the Noometry Index (29.2 vs 28.6), so choose on price, context window or the category you care about most.

Last verified . 47 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

GPT-4o OpenAI

28.6

Rank #324 Confirmed

Summary

  • They share 47 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 4 categories and GPT-4o in 6 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-4o leads 28.8 to 18.7.
  • The biggest single-benchmark swing is GeoBench: 34% for Claude 3.5 Haiku and 71% for GPT-4o.

Side by side

Claude 3.5 Haiku and GPT-4o specifications
Claude 3.5 HaikuGPT-4o
ProviderAnthropicOpenAI
Noometry Index29.228.6
Released2024-10-222024-05-13
WeightsProprietaryProprietary
Context window—128K
Max output—16K
Input $ / M tokens—$2.50
Output $ / M tokens—$10
Results tracked4972

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Haiku leads

Claude 3.5 Haiku: 32.9 (#265), GPT-4o: 24.8 (#328)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
Aider Polyglot28%45.3%
WeirdML30.7%25.1%
BigCodeBench Instruct46.1%51.1%
LiveBench Coding51.4%51.4%
LMArena Coding12861297
BigCodeBench Complete59%61.1%
CadEval32%26%
SWE-bench Verified—31%
SWE-bench Verified (bash only)—21.6%
SciCode27.4%—
GSO—0%
HumanEval+—87.2%
MBPP+—72.2%

Agentic & Tool Use Claude 3.5 Haiku leads

Claude 3.5 Haiku: 28.0 (#95), GPT-4o: 21.0 (#141)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
BALROG19.3%32.3%
GDPval—9.9%
TheAgentCompany—8.6%
Cybench—12.5%
LMArena Search—1006
METR Time Horizons—40.8%

Reasoning Claude 3.5 Haiku leads

Claude 3.5 Haiku: 17.7 (#290), GPT-4o: 9.4 (#343)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
CritPt0%0%
LiveBench Reasoning28.1%55.8%
LMArena Hard Prompts12511281
DTBench56.7%64.5%
LiveBench Data Analysis48.5%60.9%
Epoch Capabilities Index127.15128.97
LiveBench43.5%55.3%
ARC-AGI-2—0%
SimpleBench—17.8%
ARC-AGI-1—4.5%
Chess Puzzles—13%
EnigmaEval—0.8%
LMCA—16.6%
ForecastBench—57.7

Math Claude 3.5 Haiku leads

Claude 3.5 Haiku: 14.7 (#300), GPT-4o: 10.6 (#312)

Math benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
OTIS Mock AIME 2024-20254.3%6.4%
Omni-MATH22.4%29.3%
LiveBench Math35.5%49.5%
LMArena Math12441285
MATH Level 546.4%53.3%
FrontierMath (Feb 2025 set)0.3%0.3%
FrontierMath (Tiers 1-3)—0.4%

Knowledge GPT-4o leads

Claude 3.5 Haiku: 18.7 (#281), GPT-4o: 28.8 (#242)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
GPQA Diamond38.1%49.2%
MMLU-Pro60.5%71.3%
Confabulations36.7%15.3%
GPQA (HELM)36.3%52%
LMArena Expert12081250
MMLU74.3%88.1%
Humanity's Last Exam—2.7%
SimpleQA Verified—26%
Vectara Hallucination Rate—9.6%

Multimodal GPT-4o leads

Claude 3.5 Haiku: 26.8 (#117), GPT-4o: 34.5 (#91)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
LMArena Vision10921137
GeoBench34%71%
Video-MME—71.9%
VPCT—40%
ScienceQA—88.5%

Multilingual GPT-4o leads

Claude 3.5 Haiku: 40.0 (#218), GPT-4o: 43.2 (#186)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
LMArena Non-English12381283
LMArena Chinese12291277
LMArena French12641304
LMArena German12371282
LMArena Japanese11751257
LMArena Korean11731234
LMArena Russian12531286
LMArena Spanish12611292

Instruction Following GPT-4o leads

Claude 3.5 Haiku: 62.9 (#234), GPT-4o: 66.6 (#207)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
LiveBench Instruction Following61.9%68.6%
IFEval79.2%81.7%
LMArena Instruction Following12411278

Long Context GPT-4o leads

Claude 3.5 Haiku: 38.3 (#200), GPT-4o: 39.4 (#179)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
LMArena Longer Query12611289
Fiction.LiveBench—66.7%

Writing & Preference GPT-4o leads

Claude 3.5 Haiku: 42.7 (#234), GPT-4o: 52.6 (#166)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGPT-4o
LMArena Text12551300
LMArena Creative Writing12331292
Short-Story Creative Writing73.5%81.8%
WildBench76%82.8%
LMArena Multi-Turn12651302
LiveBench Language35.4%47.6%
EQ-Bench Creative Writing1146—

Frequently asked questions

Is Claude 3.5 Haiku better than GPT-4o?

Claude 3.5 Haiku and GPT-4o score almost the same on the Noometry Index (29.2 vs 28.6), so choose on price, context window or the category you care about most.

Is Claude 3.5 Haiku or GPT-4o better for coding?

Claude 3.5 Haiku scores higher on coding benchmarks: 32.9 versus 24.8 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and GPT-4o share?

47 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and GPT-4o has 72.

Related comparisons

Go deeper