Model comparison

Claude Sonnet 5.5 vs Muse Spark 1.2

Claude Sonnet 5.5 is the stronger model overall, scoring 61.9 to 50.3 on the Noometry Index. Muse Spark 1.2 costs 2.0× less per token, which makes it the better buy when Claude Sonnet 5.5's lead doesn't matter for your workload.

Last verified . 22 shared benchmarks.

Claude Sonnet 5.5 Anthropic

61.9

Rank #10 Confirmed

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Claude Sonnet 5.5 scores higher in 8 categories and Muse Spark 1.2 in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 5.5 leads 87.9 to 46.4.
  • The biggest single-benchmark swing is ProofBench: 100% for Claude Sonnet 5.5 and 43% for Muse Spark 1.2.
  • Muse Spark 1.2 is cheaper at $1.25 / $4.25 per million input/output tokens, against $2 / $10 for Claude Sonnet 5.5.
  • Muse Spark 1.2 accepts more context: 1.05M tokens versus 1M.

Side by side

Claude Sonnet 5.5 and Muse Spark 1.2 specifications
Claude Sonnet 5.5Muse Spark 1.2
ProviderAnthropicMeta
Noometry Index61.950.3
Released2026-09-282026-08-05
WeightsProprietaryProprietary
Context window1M1.05M
Max output128K131K
Input $ / M tokens$2$1.25
Output $ / M tokens$10$4.25
Results tracked3231

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 67.3 (#6), Muse Spark 1.2: 49.2 (#51)

Coding benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
LMArena WebDev17741533
FrontierSWE61.9%12%
SciCode61%56.4%
LMArena Coding15131495
DeepSWE—54.9%
FrontierCode52.1%—
CursorBench55.5%—
WeirdML—60.3%
ALE-Bench1,819—

Agentic & Tool Use Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 45.0 (#16), Muse Spark 1.2: 29.4 (#87)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
APEX-Agents75.5%36.4%
GDP.pdf—16%

Reasoning Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 54.0 (#28), Muse Spark 1.2: 51.3 (#34)

Reasoning benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
NYT Connections (extended)80.5%79.2%
CritPt31.4%17.7%
LMArena Hard Prompts14951486
Epoch Capabilities Index165.03154.87
SimpleBench—74.5%
Mystery Game Puzzles65%—
DTBench—94.7%
LMCA—48.4%

Math Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 87.9 (#6), Muse Spark 1.2: 46.4 (#70)

Math benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
ProofBench100%43%
LMArena Math15101471
FrontierMath (Tiers 1-3)88.8%—
FrontierMath Tier 480.5%—
OTIS Mock AIME 2024-2025100%—
FrontierMath Erdős2.9%—

Knowledge Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 66.0 (#12), Muse Spark 1.2: 54.1 (#53)

Knowledge benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
SimpleQA Verified46.5%60.3%
LMArena Expert15401480
GPQA Diamond95.6%—

Multimodal Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 51.5 (#6), Muse Spark 1.2: 43.4 (#25)

Multimodal benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
LMArena Vision12891305
Furniture Assembly75%—

Multilingual Muse Spark 1.2 leads

Claude Sonnet 5.5: 55.3 (#30), Muse Spark 1.2: 57.1 (#11)

Multilingual benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
LMArena Non-English14521478
LMArena Chinese15221511
LMArena Russian14511487
LMArena French—1513
LMArena Spanish—1498

Instruction Following Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 78.3 (#11), Muse Spark 1.2: 76.7 (#36)

Instruction Following benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
LMArena Instruction Following14951461

Long Context Too close to call

Claude Sonnet 5.5: 45.9 (#28), Muse Spark 1.2: 45.2 (#48)

Long Context benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
LMArena Longer Query14981475

Writing & Preference Muse Spark 1.2 leads

Claude Sonnet 5.5: 66.0 (#40), Muse Spark 1.2: 72.3 (#14)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 5.5Muse Spark 1.2
LMArena Text14711482
LMArena Creative Writing14651449
LMArena Multi-Turn14741494
EQ-Bench Creative Writing—1840

Frequently asked questions

Is Claude Sonnet 5.5 better than Muse Spark 1.2?

Claude Sonnet 5.5 is the stronger model overall, scoring 61.9 to 50.3 on the Noometry Index. Muse Spark 1.2 costs 2.0× less per token, which makes it the better buy when Claude Sonnet 5.5's lead doesn't matter for your workload.

Which is cheaper, Claude Sonnet 5.5 or Muse Spark 1.2?

Muse Spark 1.2 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; Claude Sonnet 5.5 lists at $2 and $10.

Is Claude Sonnet 5.5 or Muse Spark 1.2 better for coding?

Claude Sonnet 5.5 scores higher on coding benchmarks: 67.3 versus 49.2 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.2 does, with 1.05M tokens against 1M.

How many benchmarks do Claude Sonnet 5.5 and Muse Spark 1.2 share?

22 benchmarks have published results for both models. Claude Sonnet 5.5 has 32 scored results on Noometry and Muse Spark 1.2 has 31.

Related comparisons

Go deeper