Model comparison

Claude 3.5 Haiku vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 29.2 on the Noometry Index.

Last verified . 25 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Grok 4.6 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.6 leads 67.0 to 14.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.3% for Claude 3.5 Haiku and 99.2% for Grok 4.6.

Side by side

Claude 3.5 Haiku and Grok 4.6 specifications
Claude 3.5 HaikuGrok 4.6
ProviderAnthropicxAI
Noometry Index29.256.9
Released2024-10-222026-08-12
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked4949

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Claude 3.5 Haiku: 32.9 (#265), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
SciCode27.4%56.5%
WeirdML30.7%67.3%
LMArena Coding12861465
DeepSWE—67.5%
FrontierCode—48%
Aider Polyglot28%—
CursorBench—41.4%
LMArena WebDev—1617
FrontierSWE—25.3%
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—
ALE-Bench—1,508

Agentic & Tool Use Grok 4.6 leads

Claude 3.5 Haiku: 28.0 (#95), Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
APEX-Agents—65.3%
BALROG19.3%—
GDP.pdf—17.2%
Vending-Bench 2—9,047

Reasoning Grok 4.6 leads

Claude 3.5 Haiku: 17.7 (#290), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
CritPt0%19.7%
LMArena Hard Prompts12511447
DTBench56.7%97.3%
Epoch Capabilities Index127.15156.44
ARC-AGI-2—67.1%
SimpleBench—75.9%
NYT Connections (extended)—80%
ARC-AGI-1—87.5%
Chess Puzzles—40%
EBR-Bench—30.5%
LiveBench Reasoning28.1%—
Mystery Game Puzzles—34%
LiveBench Data Analysis48.5%—
LMCA—48.5%
LiveBench43.5%—

Math Grok 4.6 leads

Claude 3.5 Haiku: 14.7 (#300), Grok 4.6: 67.0 (#24)

Math benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
OTIS Mock AIME 2024-20254.3%99.2%
LMArena Math12441423
FrontierMath (Tiers 1-3)—66%
FrontierMath Tier 4—31.7%
ProofBench—51%
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Grok 4.6 leads

Claude 3.5 Haiku: 18.7 (#281), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
GPQA Diamond38.1%94%
LMArena Expert12081467
SimpleQA Verified—49.3%
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Grok 4.6 leads

Claude 3.5 Haiku: 26.8 (#117), Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
LMArena Vision10921263
GeoBench34%—
Blueprint-Bench 2—33.2%
Furniture Assembly—40%
LMArena Document—1452

Multilingual Grok 4.6 leads

Claude 3.5 Haiku: 40.0 (#218), Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
LMArena Non-English12381420
LMArena Chinese12291480
LMArena French12641461
LMArena German12371431
LMArena Japanese11751376
LMArena Korean11731397
LMArena Russian12531422
LMArena Spanish12611404

Instruction Following Grok 4.6 leads

Claude 3.5 Haiku: 62.9 (#234), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
LMArena Instruction Following12411431
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Grok 4.6 leads

Claude 3.5 Haiku: 38.3 (#200), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
LMArena Longer Query12611454

Writing & Preference Grok 4.6 leads

Claude 3.5 Haiku: 42.7 (#234), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuGrok 4.6
LMArena Text12551428
LMArena Creative Writing12331428
LMArena Multi-Turn12651425
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Grok 4.6 share?

25 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper