Model comparison

Grok 4.7 vs Muse Spark

Grok 4.7 is the stronger model overall, scoring 53.1 to 50.6 on the Noometry Index.

Last verified . 21 shared benchmarks.

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Grok 4.7 scores higher in 4 categories and Muse Spark in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.7 leads 49.1 to 35.9.
  • The biggest single-benchmark swing is ProofBench: 34% for Grok 4.7 and 17% for Muse Spark.

Side by side

Grok 4.7 and Muse Spark specifications
Grok 4.7Muse Spark
ProviderxAIMeta
Noometry Index53.150.6
Released2026-09-212026-04-08
WeightsProprietaryProprietary
Context window500K—
Max output500K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.7 leads

Grok 4.7: 58.0 (#18), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGrok 4.7Muse Spark
SciCode57.8%51.5%
LMArena Coding14271481
FrontierCode47.6%—
CursorBench46.3%—
LMArena WebDev1639—
FrontierSWE29.5%—

Agentic & Tool Use Not comparable

Grok 4.7: 36.7 (#37), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.7Muse Spark
APEX-Agents54.6%—
GDP.pdf22.8%—
Vending-Bench 210,537—

Reasoning Grok 4.7 leads

Grok 4.7: 49.1 (#40), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGrok 4.7Muse Spark
CritPt18%11.3%
LMArena Hard Prompts14131474
Epoch Capabilities Index153.53152.04
NYT Connections (extended)76.8%—
Chess Puzzles38%—
Mystery Game Puzzles29%—
DTBench96%—
LMCA49.4%—

Math Grok 4.7 leads

Grok 4.7: 57.8 (#39), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGrok 4.7Muse Spark
OTIS Mock AIME 2024-202598.1%88.9%
ProofBench34%17%
LMArena Math14071455
FrontierMath (Tiers 1-3)53%—
FrontierMath Tier 417.1%—
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

Grok 4.7: 62.8 (#22), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGrok 4.7Muse Spark
GPQA Diamond92.7%89.8%
LMArena Expert14221457
Humanity's Last Exam—40.6%
SimpleQA Verified56%—

Multimodal Muse Spark leads

Grok 4.7: 35.5 (#87), Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGrok 4.7Muse Spark
LMArena Vision12281306
Blueprint-Bench 232.5%—
Furniture Assembly20.8%—
LMArena Document—1444

Multilingual Muse Spark leads

Grok 4.7: 50.8 (#116), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGrok 4.7Muse Spark
LMArena Non-English13891464
LMArena Chinese14551509
LMArena French14551497
LMArena Russian13971466
LMArena Spanish14001472
LMArena German—1497
LMArena Korean—1459

Instruction Following Muse Spark leads

Grok 4.7: 74.1 (#105), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGrok 4.7Muse Spark
LMArena Instruction Following14041442

Long Context Muse Spark leads

Grok 4.7: 43.1 (#104), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGrok 4.7Muse Spark
LMArena Longer Query14131451

Writing & Preference Grok 4.7 leads

Grok 4.7: 70.0 (#24), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGrok 4.7Muse Spark
LMArena Text13991474
LMArena Creative Writing13911459
LMArena Multi-Turn13931477
EQ-Bench Creative Writing2007—

Frequently asked questions

Is Grok 4.7 better than Muse Spark?

Grok 4.7 is the stronger model overall, scoring 53.1 to 50.6 on the Noometry Index.

Is Grok 4.7 or Muse Spark better for coding?

Grok 4.7 scores higher on coding benchmarks: 58.0 versus 46.2 in the Noometry coding category.

How many benchmarks do Grok 4.7 and Muse Spark share?

21 benchmarks have published results for both models. Grok 4.7 has 39 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper