Model comparison

GPT-5.1 vs Muse Spark 1.1

GPT-5.1 and Muse Spark 1.1 score almost the same on the Noometry Index (49.0 vs 49.9), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Summary

  • They share 27 benchmarks with published results for both. GPT-5.1 scores higher in 5 categories and Muse Spark 1.1 in 5 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Muse Spark 1.1 leads 73.4 to 64.5.
  • The biggest single-benchmark swing is SciCode: 43.3% for GPT-5.1 and 58.8% for Muse Spark 1.1.
  • Muse Spark 1.1 is cheaper at $1.25 / $4.25 per million input/output tokens, against $1.25 / $10 for GPT-5.1.
  • Muse Spark 1.1 accepts more context: 1.05M tokens versus 400K.

Side by side

GPT-5.1 and Muse Spark 1.1 specifications
GPT-5.1Muse Spark 1.1
ProviderOpenAIMeta
Noometry Index49.049.9
Released2025-11-132026-04-08
WeightsProprietaryProprietary
Context window400K1.05M
Max output128K131K
Input $ / M tokens$1.25$1.25
Output $ / M tokens$10$4.25
Results tracked6337

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.1 leads

GPT-5.1: 46.4 (#66), Muse Spark 1.1: 51.3 (#40)

Coding benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
LMArena WebDev13951542
SciCode43.3%58.8%
LMArena Coding14541498
SWE-bench Verified68%—
DeepSWE—53.3%
SWE-bench Verified (bash only)66%—
GSO13.7%—
WeirdML60.8%—
LiveBench Coding72.5%—
ALE-Bench1,192—

Agentic & Tool Use GPT-5.1 leads

GPT-5.1: 32.7 (#60), Muse Spark 1.1: 30.8 (#73)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
Vending-Bench 21,4736,520
Terminal-Bench47.6%—
APEX-Agents—31.8%
τ²-bench Banking—40.5%
DeepResearch Bench42.8%—
GBAEval—7.9%
GDP.pdf—15%
LMArena Search1199—

Reasoning Muse Spark 1.1 leads

GPT-5.1: 39.8 (#58), Muse Spark 1.1: 47.1 (#44)

Reasoning benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
CritPt4.9%15.1%
LMArena Hard Prompts14571486
DTBench90.1%94.4%
LMCA43.9%49.9%
Epoch Capabilities Index149.64154.21
ARC-AGI-217.6%—
SimpleBench53.2%—
NYT Connections (extended)—84.9%
ARC-AGI-172.8%—
Chess Puzzles32%—
EnigmaEval11.2%—
LiveBench Reasoning95.8%—
Mystery Game Puzzles19%—
LiveBench Data Analysis72.1%—
Surface Evolver Bench—52.5%
ForecastBench58.1—
LiveBench78.8%—

Math GPT-5.1 leads

GPT-5.1: 52.2 (#51), Muse Spark 1.1: 45.5 (#76)

Math benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
LMArena Math14471483
OTIS Mock AIME 2024-202588.6%—
ProofBench—39%
Omni-MATH46.4%—
LiveBench Math94.5%—
FrontierMath (Feb 2025 set)31%—
FrontierMath Tier 4 (v1)12.5%—

Knowledge Muse Spark 1.1 leads

GPT-5.1: 50.6 (#71), Muse Spark 1.1: 53.1 (#59)

Knowledge benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
SimpleQA Verified48%57.8%
LMArena Expert14701478
GPQA Diamond87.6%—
Humanity's Last Exam23.7%—
MMLU-Pro57.9%—
Vectara Hallucination Rate10.9%—
GPQA (HELM)44.2%—

Multimodal GPT-5.1 leads

GPT-5.1: 44.8 (#19), Muse Spark 1.1: 42.6 (#29)

Multimodal benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
LMArena Vision12501293
LMArena Document14031465
VPCT58.7%—

Multilingual Muse Spark 1.1 leads

GPT-5.1: 53.8 (#56), Muse Spark 1.1: 56.7 (#17)

Multilingual benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
LMArena Non-English14311472
LMArena Chinese14951518
LMArena French14501494
LMArena German14381466
LMArena Japanese14531451
LMArena Korean14011458
LMArena Russian14351483
LMArena Spanish14331464

Instruction Following GPT-5.1 leads

GPT-5.1: 83.9 (#1), Muse Spark 1.1: 76.5 (#39)

Instruction Following benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
LMArena Instruction Following14431457
LiveBench Instruction Following93.3%—
IFEval93.5%—

Long Context GPT-5.1 leads

GPT-5.1: 47.6 (#14), Muse Spark 1.1: 44.8 (#58)

Long Context benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
LMArena Longer Query14471462
CL-bench23.7%—
CL-bench Life17.3%—

Writing & Preference Muse Spark 1.1 leads

GPT-5.1: 64.5 (#55), Muse Spark 1.1: 73.4 (#11)

Writing & Preference benchmarks
BenchmarkGPT-5.1Muse Spark 1.1
LMArena Text14431479
LMArena Creative Writing14271437
LMArena Multi-Turn14501485
EQ-Bench Creative Writing—1927
WildBench86.3%—
EQ-Bench 4—1260
LiveBench Language80.2%—

Frequently asked questions

Is GPT-5.1 better than Muse Spark 1.1?

GPT-5.1 and Muse Spark 1.1 score almost the same on the Noometry Index (49.0 vs 49.9), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.1 or Muse Spark 1.1?

Muse Spark 1.1 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; GPT-5.1 lists at $1.25 and $10.

Is GPT-5.1 or Muse Spark 1.1 better for coding?

Muse Spark 1.1 scores higher on coding benchmarks: 51.3 versus 46.4 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.1 does, with 1.05M tokens against 400K.

How many benchmarks do GPT-5.1 and Muse Spark 1.1 share?

27 benchmarks have published results for both models. GPT-5.1 has 63 scored results on Noometry and Muse Spark 1.1 has 37.

Related comparisons

Go deeper