Model comparison

Grok 4.5 vs Muse Spark

Grok 4.5 is the stronger model overall, scoring 55.0 to 50.6 on the Noometry Index.

Last verified . 24 shared benchmarks.

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Grok 4.5 scores higher in 5 categories and Muse Spark in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.5 leads 56.1 to 35.9.
  • The biggest single-benchmark swing is ProofBench: 31% for Grok 4.5 and 17% for Muse Spark.

Side by side

Grok 4.5 and Muse Spark specifications
Grok 4.5Muse Spark
ProviderxAIMeta
Noometry Index55.050.6
Released2026-07-082026-04-08
WeightsProprietaryProprietary
Context window500K—
Max output500K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked5227

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.5 leads

Grok 4.5: 52.2 (#35), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGrok 4.5Muse Spark
SciCode54.1%51.5%
LMArena Coding14741481
DeepSWE53.8%—
FrontierCode42.4%—
LMArena WebDev1553—
WeirdML46.4%—
ALE-Bench1,309—

Agentic & Tool Use Not comparable

Grok 4.5: 44.4 (#17), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.5Muse Spark
APEX-Agents56.2%—
τ²-bench Banking47.9%—
PostTrainBench23.4%—
GBAEval65.4%—
GDP.pdf14%—
LMArena Search1213—
Vending-Bench 23,887—

Reasoning Grok 4.5 leads

Grok 4.5: 56.1 (#25), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGrok 4.5Muse Spark
CritPt15.4%11.3%
LMArena Hard Prompts14621474
Epoch Capabilities Index153.92152.04
ARC-AGI-252.6%—
SimpleBench70%—
Kagi LLM Benchmark83.5%—
NYT Connections (extended)79.9%—
ARC-AGI-187.2%—
Chess Puzzles36%—
DTBench96.5%—
LMCA45.2%—
Surface Evolver Bench74.4%—

Math Grok 4.5 leads

Grok 4.5: 60.9 (#35), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGrok 4.5Muse Spark
OTIS Mock AIME 2024-202597.8%88.9%
ProofBench31%17%
LMArena Math14591455
FrontierMath (Tiers 1-3)57.2%—
FrontierMath Tier 424.4%—
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

Grok 4.5: 62.3 (#24), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGrok 4.5Muse Spark
GPQA Diamond93.4%89.8%
LMArena Expert14661457
Humanity's Last Exam—40.6%
SimpleQA Verified48.3%—

Multimodal Muse Spark leads

Grok 4.5: 37.6 (#72), Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGrok 4.5Muse Spark
LMArena Vision12881306
LMArena Document14521444
Blueprint-Bench 227.3%—
Furniture Assembly22.5%—

Multilingual Muse Spark leads

Grok 4.5: 54.4 (#42), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGrok 4.5Muse Spark
LMArena Non-English14401464
LMArena Chinese14961509
LMArena French14561497
LMArena German14461497
LMArena Korean14041459
LMArena Russian14481466
LMArena Spanish14501472
LMArena Japanese1428—

Instruction Following Too close to call

Grok 4.5: 76.0 (#48), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGrok 4.5Muse Spark
LMArena Instruction Following14461442

Long Context Too close to call

Grok 4.5: 44.8 (#56), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGrok 4.5Muse Spark
LMArena Longer Query14631451

Writing & Preference Too close to call

Grok 4.5: 65.8 (#42), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGrok 4.5Muse Spark
LMArena Text14481474
LMArena Creative Writing14421459
LMArena Multi-Turn14561477
EQ-Bench Creative Writing1579—

Frequently asked questions

Is Grok 4.5 better than Muse Spark?

Grok 4.5 is the stronger model overall, scoring 55.0 to 50.6 on the Noometry Index.

Is Grok 4.5 or Muse Spark better for coding?

Grok 4.5 scores higher on coding benchmarks: 52.2 versus 46.2 in the Noometry coding category.

How many benchmarks do Grok 4.5 and Muse Spark share?

24 benchmarks have published results for both models. Grok 4.5 has 52 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper