Model comparison

GPT-5 vs Muse Spark 1.1

GPT-5 and Muse Spark 1.1 score almost the same on the Noometry Index (50.9 vs 49.9), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

GPT-5 OpenAI

50.9

Rank #45 Confirmed

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Summary

  • They share 27 benchmarks with published results for both. GPT-5 scores higher in 5 categories and Muse Spark 1.1 in 5 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where GPT-5 leads 69.5 to 44.8.
  • The biggest single-benchmark swing is ProofBench: 18% for GPT-5 and 39% for Muse Spark 1.1.
  • Muse Spark 1.1 is cheaper at $1.25 / $4.25 per million input/output tokens, against $1.25 / $10 for GPT-5.
  • Muse Spark 1.1 accepts more context: 1.05M tokens versus 400K.

Side by side

GPT-5 and Muse Spark 1.1 specifications
GPT-5Muse Spark 1.1
ProviderOpenAIMeta
Noometry Index50.949.9
Released2025-08-072026-04-08
WeightsProprietaryProprietary
Context window400K1.05M
Max output128K131K
Input $ / M tokens$1.25$1.25
Output $ / M tokens$10$4.25
Results tracked6937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.1 leads

GPT-5: 50.3 (#47), Muse Spark 1.1: 51.3 (#40)

Coding benchmarks
BenchmarkGPT-5Muse Spark 1.1
LMArena WebDev14181542
SciCode42.9%58.8%
LMArena Coding14361498
SWE-bench Verified73.6%—
DeepSWE—53.3%
SWE-bench Verified (bash only)65%—
Aider Polyglot88%—
GSO6.9%—
WeirdML60.7%—
ALE-Bench1,162—
AlgoTune1.67—

Agentic & Tool Use GPT-5 leads

GPT-5: 33.1 (#56), Muse Spark 1.1: 30.8 (#73)

Agentic & Tool Use benchmarks
BenchmarkGPT-5Muse Spark 1.1
Terminal-Bench49.6%—
APEX-Agents—31.8%
GDPval34.8%—
Remote Labor Index1.7%—
τ²-bench Banking—40.5%
DeepResearch Bench49.6%—
BALROG32.8%—
GBAEval—7.9%
GDP.pdf—15%
LMArena Search1133—
METR Time Horizons69.6%—
Vending-Bench 2—6,520

Reasoning Muse Spark 1.1 leads

GPT-5: 38.3 (#64), Muse Spark 1.1: 47.1 (#44)

Reasoning benchmarks
BenchmarkGPT-5Muse Spark 1.1
CritPt12.6%15.1%
LMArena Hard Prompts14161486
DTBench90.7%94.4%
LMCA40%49.9%
Epoch Capabilities Index150154.21
ARC-AGI-29.9%—
SimpleBench56.7%—
Kagi LLM Benchmark72.7%—
NYT Connections (extended)—84.9%
ARC-AGI-165.7%—
Chess Puzzles37%—
EnigmaEval10.5%—
EBR-Bench12.7%—
Mystery Game Puzzles23%—
Surface Evolver Bench—52.5%
ForecastBench61.4—

Math GPT-5 leads

GPT-5: 55.0 (#44), Muse Spark 1.1: 45.5 (#76)

Math benchmarks
BenchmarkGPT-5Muse Spark 1.1
ProofBench18%39%
LMArena Math14071483
FrontierMath (Tiers 1-3)55.4%—
FrontierMath Tier 422%—
OTIS Mock AIME 2024-202591.4%—
Omni-MATH64.7%—
MATH Level 598.1%—
FrontierMath (Feb 2025 set)32.4%—
FrontierMath Tier 4 (v1)12.5%—

Knowledge GPT-5 leads

GPT-5: 56.6 (#43), Muse Spark 1.1: 53.1 (#59)

Knowledge benchmarks
BenchmarkGPT-5Muse Spark 1.1
SimpleQA Verified50.1%57.8%
LMArena Expert14191478
GPQA Diamond86.2%—
Humanity's Last Exam25.3%—
MMLU-Pro86.3%—
Confabulations10.3%—
Vectara Hallucination Rate14.7%—
GPQA (HELM)79.2%—

Multimodal GPT-5 leads

GPT-5: 46.8 (#13), Muse Spark 1.1: 42.6 (#29)

Multimodal benchmarks
BenchmarkGPT-5Muse Spark 1.1
LMArena Vision12321293
GeoBench81%—
VPCT66%—
LMArena Document—1465

Multilingual Muse Spark 1.1 leads

GPT-5: 51.4 (#110), Muse Spark 1.1: 56.7 (#17)

Multilingual benchmarks
BenchmarkGPT-5Muse Spark 1.1
LMArena Non-English13971472
LMArena Chinese14221518
LMArena French14101494
LMArena German14161466
LMArena Japanese14091451
LMArena Korean13601458
LMArena Russian14061483
LMArena Spanish13991464

Instruction Following Muse Spark 1.1 leads

GPT-5: 73.8 (#113), Muse Spark 1.1: 76.5 (#39)

Instruction Following benchmarks
BenchmarkGPT-5Muse Spark 1.1
LMArena Instruction Following13881457
IFEval87.5%—

Long Context GPT-5 leads

GPT-5: 69.5 (#2), Muse Spark 1.1: 44.8 (#58)

Long Context benchmarks
BenchmarkGPT-5Muse Spark 1.1
LMArena Longer Query13991462
Fiction.LiveBench97.2%—

Writing & Preference Muse Spark 1.1 leads

GPT-5: 63.4 (#65), Muse Spark 1.1: 73.4 (#11)

Writing & Preference benchmarks
BenchmarkGPT-5Muse Spark 1.1
LMArena Text14061479
LMArena Creative Writing13651437
EQ-Bench Creative Writing16271927
LMArena Multi-Turn14261485
Short-Story Creative Writing86%—
WildBench85.7%—
EQ-Bench 4—1260

Frequently asked questions

Is GPT-5 better than Muse Spark 1.1?

GPT-5 and Muse Spark 1.1 score almost the same on the Noometry Index (50.9 vs 49.9), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 or Muse Spark 1.1?

Muse Spark 1.1 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; GPT-5 lists at $1.25 and $10.

Is GPT-5 or Muse Spark 1.1 better for coding?

Muse Spark 1.1 scores higher on coding benchmarks: 51.3 versus 50.3 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.1 does, with 1.05M tokens against 400K.

How many benchmarks do GPT-5 and Muse Spark 1.1 share?

27 benchmarks have published results for both models. GPT-5 has 69 scored results on Noometry and Muse Spark 1.1 has 37.

Related comparisons

Go deeper