Model comparison

Grok 4.3 vs Muse Spark 1.3

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 43.8 on the Noometry Index.

Last verified . 31 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Grok 4.3 scores higher in 1 category and Muse Spark 1.3 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Muse Spark 1.3 leads 73.1 to 46.0.
  • The biggest single-benchmark swing is ProofBench: 11% for Grok 4.3 and 58% for Muse Spark 1.3.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $1.25 / $4.25 for Muse Spark 1.3.
  • Muse Spark 1.3 accepts more context: 1.05M tokens versus 1M.

Side by side

Grok 4.3 and Muse Spark 1.3 specifications
Grok 4.3Muse Spark 1.3
ProviderxAIMeta
Noometry Index43.854.8
Released2026-04-172026-09-02
WeightsProprietaryProprietary
Context window1M1.05M
Max output30K131K
Input $ / M tokens$1.25$1.25
Output $ / M tokens$2.50$4.25
Results tracked4037

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.3 leads

Grok 4.3: 41.6 (#121), Muse Spark 1.3: 56.6 (#21)

Coding benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
LMArena WebDev13571657
SciCode47.3%59.7%
LMArena Coding14151514
CursorBench—41.6%
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Muse Spark 1.3 leads

Grok 4.3: 27.7 (#99), Muse Spark 1.3: 38.6 (#30)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
GDP.pdf8%27.6%
APEX-Agents—57.8%
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Muse Spark 1.3 leads

Grok 4.3: 35.9 (#68), Muse Spark 1.3: 54.0 (#27)

Reasoning benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
NYT Connections (extended)55.2%85.1%
CritPt8%26%
Chess Puzzles25%38%
LMArena Hard Prompts13961503
DTBench90.7%96.5%
LMCA38.3%53.9%
Epoch Capabilities Index149.16156.75
Mystery Game Puzzles—25%
Bench to the Future 3—0.14
ForecastBench60.3—

Math Muse Spark 1.3 leads

Grok 4.3: 46.0 (#74), Muse Spark 1.3: 73.1 (#21)

Math benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
FrontierMath (Tiers 1-3)42.8%74.4%
FrontierMath Tier 414.6%46.3%
OTIS Mock AIME 2024-202593.3%99.2%
ProofBench11%58%
LMArena Math13881494

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Muse Spark 1.3: 42.6 (#95)

Knowledge benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
LMArena Expert13851516
GPQA Diamond88.8%—
SimpleQA Verified33.2%—

Multimodal Muse Spark 1.3 leads

Grok 4.3: 31.6 (#104), Muse Spark 1.3: 43.7 (#22)

Multimodal benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
LMArena Vision12291309
Blueprint-Bench 20%—
LMArena Document—1471

Multilingual Muse Spark 1.3 leads

Grok 4.3: 50.5 (#120), Muse Spark 1.3: 57.4 (#8)

Multilingual benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
LMArena Non-English13851481
LMArena Chinese14221529
LMArena French14121524
LMArena German13951515
LMArena Japanese13791474
LMArena Korean13561501
LMArena Russian13991490
LMArena Spanish13981490

Instruction Following Muse Spark 1.3 leads

Grok 4.3: 72.1 (#140), Muse Spark 1.3: 77.5 (#22)

Instruction Following benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
LMArena Instruction Following13661477

Long Context Muse Spark 1.3 leads

Grok 4.3: 42.5 (#123), Muse Spark 1.3: 45.6 (#32)

Long Context benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
LMArena Longer Query13931488

Writing & Preference Muse Spark 1.3 leads

Grok 4.3: 58.5 (#118), Muse Spark 1.3: 73.6 (#9)

Writing & Preference benchmarks
BenchmarkGrok 4.3Muse Spark 1.3
LMArena Text13971490
LMArena Creative Writing13801455
LMArena Multi-Turn14061482
EQ-Bench Creative Writing—1906
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Muse Spark 1.3?

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 43.8 on the Noometry Index.

Which is cheaper, Grok 4.3 or Muse Spark 1.3?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Muse Spark 1.3 lists at $1.25 and $4.25.

Is Grok 4.3 or Muse Spark 1.3 better for coding?

Muse Spark 1.3 scores higher on coding benchmarks: 56.6 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.3 does, with 1.05M tokens against 1M.

How many benchmarks do Grok 4.3 and Muse Spark 1.3 share?

31 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Muse Spark 1.3 has 37.

Related comparisons

Go deeper