Model comparison

Grok 4.6 vs Muse Spark 1.2

Grok 4.6 is the stronger model overall, scoring 56.9 to 50.3 on the Noometry Index. Muse Spark 1.2 costs 1.5× less per token, which makes it the better buy when Grok 4.6's lead doesn't matter for your workload.

Last verified . 30 shared benchmarks.

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Grok 4.6 scores higher in 6 categories and Muse Spark 1.2 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.6 leads 67.0 to 46.4.
  • The biggest single-benchmark swing is APEX-Agents: 65.3% for Grok 4.6 and 36.4% for Muse Spark 1.2.
  • Muse Spark 1.2 is cheaper at $1.25 / $4.25 per million input/output tokens, against $2 / $6 for Grok 4.6.
  • Muse Spark 1.2 accepts more context: 1.05M tokens versus 500K.

Side by side

Grok 4.6 and Muse Spark 1.2 specifications
Grok 4.6Muse Spark 1.2
ProviderxAIMeta
Noometry Index56.950.3
Released2026-08-122026-08-05
WeightsProprietaryProprietary
Context window500K1.05M
Max output500K131K
Input $ / M tokens$2$1.25
Output $ / M tokens$6$4.25
Results tracked4931

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Grok 4.6: 58.5 (#16), Muse Spark 1.2: 49.2 (#51)

Coding benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
DeepSWE67.5%54.9%
LMArena WebDev16171533
FrontierSWE25.3%12%
SciCode56.5%56.4%
WeirdML67.3%60.3%
LMArena Coding14651495
FrontierCode48%—
CursorBench41.4%—
ALE-Bench1,508—

Agentic & Tool Use Grok 4.6 leads

Grok 4.6: 39.4 (#27), Muse Spark 1.2: 29.4 (#87)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
APEX-Agents65.3%36.4%
GDP.pdf17.2%16%
Vending-Bench 29,047—

Reasoning Grok 4.6 leads

Grok 4.6: 61.4 (#20), Muse Spark 1.2: 51.3 (#34)

Reasoning benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
SimpleBench75.9%74.5%
NYT Connections (extended)80%79.2%
CritPt19.7%17.7%
LMArena Hard Prompts14471486
DTBench97.3%94.7%
LMCA48.5%48.4%
Epoch Capabilities Index156.44154.87
ARC-AGI-267.1%—
ARC-AGI-187.5%—
Chess Puzzles40%—
EBR-Bench30.5%—
Mystery Game Puzzles34%—

Math Grok 4.6 leads

Grok 4.6: 67.0 (#24), Muse Spark 1.2: 46.4 (#70)

Math benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
ProofBench51%43%
LMArena Math14231471
FrontierMath (Tiers 1-3)66%—
FrontierMath Tier 431.7%—
OTIS Mock AIME 2024-202599.2%—

Knowledge Grok 4.6 leads

Grok 4.6: 63.3 (#20), Muse Spark 1.2: 54.1 (#53)

Knowledge benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
SimpleQA Verified49.3%60.3%
LMArena Expert14671480
GPQA Diamond94%—

Multimodal Too close to call

Grok 4.6: 43.6 (#23), Muse Spark 1.2: 43.4 (#25)

Multimodal benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
LMArena Vision12631305
Blueprint-Bench 233.2%—
Furniture Assembly40%—
LMArena Document1452—

Multilingual Muse Spark 1.2 leads

Grok 4.6: 53.0 (#74), Muse Spark 1.2: 57.1 (#11)

Multilingual benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
LMArena Non-English14201478
LMArena Chinese14801511
LMArena French14611513
LMArena Russian14221487
LMArena Spanish14041498
LMArena German1431—
LMArena Japanese1376—
LMArena Korean1397—

Instruction Following Muse Spark 1.2 leads

Grok 4.6: 75.4 (#63), Muse Spark 1.2: 76.7 (#36)

Instruction Following benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
LMArena Instruction Following14311461

Long Context Too close to call

Grok 4.6: 44.5 (#66), Muse Spark 1.2: 45.2 (#48)

Long Context benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
LMArena Longer Query14541475

Writing & Preference Muse Spark 1.2 leads

Grok 4.6: 62.3 (#80), Muse Spark 1.2: 72.3 (#14)

Writing & Preference benchmarks
BenchmarkGrok 4.6Muse Spark 1.2
LMArena Text14281482
LMArena Creative Writing14281449
LMArena Multi-Turn14251494
EQ-Bench Creative Writing—1840

Frequently asked questions

Is Grok 4.6 better than Muse Spark 1.2?

Grok 4.6 is the stronger model overall, scoring 56.9 to 50.3 on the Noometry Index. Muse Spark 1.2 costs 1.5× less per token, which makes it the better buy when Grok 4.6's lead doesn't matter for your workload.

Which is cheaper, Grok 4.6 or Muse Spark 1.2?

Muse Spark 1.2 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; Grok 4.6 lists at $2 and $6.

Is Grok 4.6 or Muse Spark 1.2 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 49.2 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.2 does, with 1.05M tokens against 500K.

How many benchmarks do Grok 4.6 and Muse Spark 1.2 share?

30 benchmarks have published results for both models. Grok 4.6 has 49 scored results on Noometry and Muse Spark 1.2 has 31.

Related comparisons

Go deeper