Model comparison

GPT-5 vs Muse Spark 1.2

GPT-5 and Muse Spark 1.2 score almost the same on the Noometry Index (50.9 vs 50.3), so choose on price, context window or the category you care about most.

Last verified . 26 shared benchmarks.

GPT-5 OpenAI

50.9

Rank #45 Confirmed

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Summary

  • They share 26 benchmarks with published results for both. GPT-5 scores higher in 6 categories and Muse Spark 1.2 in 4 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where GPT-5 leads 69.5 to 45.2.
  • The biggest single-benchmark swing is ProofBench: 18% for GPT-5 and 43% for Muse Spark 1.2.
  • Muse Spark 1.2 is cheaper at $1.25 / $4.25 per million input/output tokens, against $1.25 / $10 for GPT-5.
  • Muse Spark 1.2 accepts more context: 1.05M tokens versus 400K.

Side by side

GPT-5 and Muse Spark 1.2 specifications
GPT-5Muse Spark 1.2
ProviderOpenAIMeta
Noometry Index50.950.3
Released2025-08-072026-08-05
WeightsProprietaryProprietary
Context window400K1.05M
Max output128K131K
Input $ / M tokens$1.25$1.25
Output $ / M tokens$10$4.25
Results tracked6931

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5 leads

GPT-5: 50.3 (#47), Muse Spark 1.2: 49.2 (#51)

Coding benchmarks
BenchmarkGPT-5Muse Spark 1.2
LMArena WebDev14181533
SciCode42.9%56.4%
WeirdML60.7%60.3%
LMArena Coding14361495
SWE-bench Verified73.6%—
DeepSWE—54.9%
SWE-bench Verified (bash only)65%—
Aider Polyglot88%—
FrontierSWE—12%
GSO6.9%—
ALE-Bench1,162—
AlgoTune1.67—

Agentic & Tool Use GPT-5 leads

GPT-5: 33.1 (#56), Muse Spark 1.2: 29.4 (#87)

Agentic & Tool Use benchmarks
BenchmarkGPT-5Muse Spark 1.2
Terminal-Bench49.6%—
APEX-Agents—36.4%
GDPval34.8%—
Remote Labor Index1.7%—
DeepResearch Bench49.6%—
BALROG32.8%—
GDP.pdf—16%
LMArena Search1133—
METR Time Horizons69.6%—

Reasoning Muse Spark 1.2 leads

GPT-5: 38.3 (#64), Muse Spark 1.2: 51.3 (#34)

Reasoning benchmarks
BenchmarkGPT-5Muse Spark 1.2
SimpleBench56.7%74.5%
CritPt12.6%17.7%
LMArena Hard Prompts14161486
DTBench90.7%94.7%
LMCA40%48.4%
Epoch Capabilities Index150154.87
ARC-AGI-29.9%—
Kagi LLM Benchmark72.7%—
NYT Connections (extended)—79.2%
ARC-AGI-165.7%—
Chess Puzzles37%—
EnigmaEval10.5%—
EBR-Bench12.7%—
Mystery Game Puzzles23%—
ForecastBench61.4—

Math GPT-5 leads

GPT-5: 55.0 (#44), Muse Spark 1.2: 46.4 (#70)

Math benchmarks
BenchmarkGPT-5Muse Spark 1.2
ProofBench18%43%
LMArena Math14071471
FrontierMath (Tiers 1-3)55.4%—
FrontierMath Tier 422%—
OTIS Mock AIME 2024-202591.4%—
Omni-MATH64.7%—
MATH Level 598.1%—
FrontierMath (Feb 2025 set)32.4%—
FrontierMath Tier 4 (v1)12.5%—

Knowledge GPT-5 leads

GPT-5: 56.6 (#43), Muse Spark 1.2: 54.1 (#53)

Knowledge benchmarks
BenchmarkGPT-5Muse Spark 1.2
SimpleQA Verified50.1%60.3%
LMArena Expert14191480
GPQA Diamond86.2%—
Humanity's Last Exam25.3%—
MMLU-Pro86.3%—
Confabulations10.3%—
Vectara Hallucination Rate14.7%—
GPQA (HELM)79.2%—

Multimodal GPT-5 leads

GPT-5: 46.8 (#13), Muse Spark 1.2: 43.4 (#25)

Multimodal benchmarks
BenchmarkGPT-5Muse Spark 1.2
LMArena Vision12321305
GeoBench81%—
VPCT66%—

Multilingual Muse Spark 1.2 leads

GPT-5: 51.4 (#110), Muse Spark 1.2: 57.1 (#11)

Multilingual benchmarks
BenchmarkGPT-5Muse Spark 1.2
LMArena Non-English13971478
LMArena Chinese14221511
LMArena French14101513
LMArena Russian14061487
LMArena Spanish13991498
LMArena German1416—
LMArena Japanese1409—
LMArena Korean1360—

Instruction Following Muse Spark 1.2 leads

GPT-5: 73.8 (#113), Muse Spark 1.2: 76.7 (#36)

Instruction Following benchmarks
BenchmarkGPT-5Muse Spark 1.2
LMArena Instruction Following13881461
IFEval87.5%—

Long Context GPT-5 leads

GPT-5: 69.5 (#2), Muse Spark 1.2: 45.2 (#48)

Long Context benchmarks
BenchmarkGPT-5Muse Spark 1.2
LMArena Longer Query13991475
Fiction.LiveBench97.2%—

Writing & Preference Muse Spark 1.2 leads

GPT-5: 63.4 (#65), Muse Spark 1.2: 72.3 (#14)

Writing & Preference benchmarks
BenchmarkGPT-5Muse Spark 1.2
LMArena Text14061482
LMArena Creative Writing13651449
EQ-Bench Creative Writing16271840
LMArena Multi-Turn14261494
Short-Story Creative Writing86%—
WildBench85.7%—

Frequently asked questions

Is GPT-5 better than Muse Spark 1.2?

GPT-5 and Muse Spark 1.2 score almost the same on the Noometry Index (50.9 vs 50.3), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 or Muse Spark 1.2?

Muse Spark 1.2 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; GPT-5 lists at $1.25 and $10.

Is GPT-5 or Muse Spark 1.2 better for coding?

GPT-5 scores higher on coding benchmarks: 50.3 versus 49.2 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.2 does, with 1.05M tokens against 400K.

How many benchmarks do GPT-5 and Muse Spark 1.2 share?

26 benchmarks have published results for both models. GPT-5 has 69 scored results on Noometry and Muse Spark 1.2 has 31.

Related comparisons

Go deeper