Model comparison

GPT-5.2 vs Muse Spark 1.3

GPT-5.2 and Muse Spark 1.3 score almost the same on the Noometry Index (54.1 vs 54.8), so choose on price, context window or the category you care about most.

Last verified . 31 shared benchmarks.

GPT-5.2 OpenAI

54.1

Rank #34 Confirmed

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Summary

  • They share 31 benchmarks with published results for both. GPT-5.2 scores higher in 3 categories and Muse Spark 1.3 in 7 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5.2 leads 59.3 to 42.6.
  • The biggest single-benchmark swing is ProofBench: 15% for GPT-5.2 and 58% for Muse Spark 1.3.
  • Muse Spark 1.3 is cheaper at $1.25 / $4.25 per million input/output tokens, against $1.75 / $14 for GPT-5.2.
  • Muse Spark 1.3 accepts more context: 1.05M tokens versus 400K.

Side by side

GPT-5.2 and Muse Spark 1.3 specifications
GPT-5.2Muse Spark 1.3
ProviderOpenAIMeta
Noometry Index54.154.8
Released2025-12-112026-09-02
WeightsProprietaryProprietary
Context window400K1.05M
Max output128K131K
Input $ / M tokens$1.75$1.25
Output $ / M tokens$14$4.25
Results tracked6737

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.3 leads

GPT-5.2: 51.6 (#37), Muse Spark 1.3: 56.6 (#21)

Coding benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
LMArena WebDev14161657
LMArena Coding14471514
SWE-bench Verified73.8%—
SWE-bench Verified (bash only)72.8%—
CursorBench—41.6%
SWE-bench Multilingual66.7%—
SciCode—59.7%
GSO27.4%—
WeirdML72.2%—
ALE-Bench1,294—
AlgoTune2.05—

Agentic & Tool Use GPT-5.2 leads

GPT-5.2: 40.2 (#24), Muse Spark 1.3: 38.6 (#30)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
Terminal-Bench64.9%—
APEX-Agents—57.8%
Berkeley Function Calling Leaderboard55.9%—
GDPval49.7%—
Remote Labor Index2.5%—
τ²-bench Airline83%—
τ²-bench Banking32.2%—
τ²-bench Retail81.6%—
τ²-bench Telecom89.7%—
DeepResearch Bench41.1%—
GDP.pdf—27.6%
LMArena Search1207—
METR Time Horizons75.3%—
Vending-Bench 23,591—

Reasoning Muse Spark 1.3 leads

GPT-5.2: 50.2 (#35), Muse Spark 1.3: 54.0 (#27)

Reasoning benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
NYT Connections (extended)83.6%85.1%
Chess Puzzles49%38%
LMArena Hard Prompts14451503
Mystery Game Puzzles23%25%
DTBench90.9%96.5%
LMCA43.9%53.9%
Epoch Capabilities Index153.45156.75
ARC-AGI-252.9%—
SimpleBench45.8%—
Kagi LLM Benchmark73.3%—
ARC-AGI-186.2%—
CritPt—26%
EnigmaEval10.4%—
EBR-Bench23%—
Bench to the Future 3—0.14
ForecastBench60.1—

Math Muse Spark 1.3 leads

GPT-5.2: 60.0 (#38), Muse Spark 1.3: 73.1 (#21)

Knowledge GPT-5.2 leads

GPT-5.2: 59.3 (#32), Muse Spark 1.3: 42.6 (#95)

Knowledge benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
LMArena Expert14451516
GPQA Diamond91.4%—
Humanity's Last Exam27.8%—
SimpleQA Verified37.1%—
Vectara Hallucination Rate8.4%—

Multimodal GPT-5.2 leads

GPT-5.2: 51.3 (#7), Muse Spark 1.3: 43.7 (#22)

Multimodal benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
LMArena Vision12681309
LMArena Document14051471
VPCT84%—
Furniture Assembly38.3%—

Multilingual Muse Spark 1.3 leads

GPT-5.2: 53.4 (#67), Muse Spark 1.3: 57.4 (#8)

Multilingual benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
LMArena Non-English14251481
LMArena Chinese14601529
LMArena French14551524
LMArena German14481515
LMArena Japanese14201474
LMArena Korean13921501
LMArena Russian14401490
LMArena Spanish14331490

Instruction Following Muse Spark 1.3 leads

GPT-5.2: 74.7 (#89), Muse Spark 1.3: 77.5 (#22)

Instruction Following benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
LMArena Instruction Following14171477

Long Context Muse Spark 1.3 leads

GPT-5.2: 44.0 (#78), Muse Spark 1.3: 45.6 (#32)

Long Context benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
LMArena Longer Query14281488
CL-bench18.2%—

Writing & Preference Muse Spark 1.3 leads

GPT-5.2: 66.8 (#32), Muse Spark 1.3: 73.6 (#9)

Writing & Preference benchmarks
BenchmarkGPT-5.2Muse Spark 1.3
LMArena Text14391490
LMArena Creative Writing14011455
EQ-Bench Creative Writing17031906
LMArena Multi-Turn14581482

Frequently asked questions

Is GPT-5.2 better than Muse Spark 1.3?

GPT-5.2 and Muse Spark 1.3 score almost the same on the Noometry Index (54.1 vs 54.8), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.2 or Muse Spark 1.3?

Muse Spark 1.3 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; GPT-5.2 lists at $1.75 and $14.

Is GPT-5.2 or Muse Spark 1.3 better for coding?

Muse Spark 1.3 scores higher on coding benchmarks: 56.6 versus 51.6 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.3 does, with 1.05M tokens against 400K.

How many benchmarks do GPT-5.2 and Muse Spark 1.3 share?

31 benchmarks have published results for both models. GPT-5.2 has 67 scored results on Noometry and Muse Spark 1.3 has 37.

Related comparisons

Go deeper