Model comparison

Claude 3.5 Sonnet vs Muse Spark 1.3

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 34.6 on the Noometry Index.

Last verified . 22 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 0 categories and Muse Spark 1.3 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Muse Spark 1.3 leads 73.1 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 99.2% for Muse Spark 1.3.

Side by side

Claude 3.5 Sonnet and Muse Spark 1.3 specifications
Claude 3.5 SonnetMuse Spark 1.3
ProviderAnthropicMeta
Noometry Index34.654.8
Released2024-06-202026-09-02
WeightsProprietaryProprietary
Context window—1.05M
Max output—131K
Input $ / M tokens—$1.25
Output $ / M tokens—$4.25
Results tracked6037

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.3 leads

Claude 3.5 Sonnet: 39.0 (#165), Muse Spark 1.3: 56.6 (#21)

Coding benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Coding13421514
Aider Polyglot51.6%—
CursorBench—41.6%
LMArena WebDev—1657
SciCode—59.7%
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Muse Spark 1.3 leads

Claude 3.5 Sonnet: 32.3 (#67), Muse Spark 1.3: 38.6 (#30)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
APEX-Agents—57.8%
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
GDP.pdf—27.6%
METR Time Horizons45.2%—

Reasoning Muse Spark 1.3 leads

Claude 3.5 Sonnet: 23.1 (#183), Muse Spark 1.3: 54.0 (#27)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Hard Prompts13051503
DTBench67.8%96.5%
Epoch Capabilities Index133.55156.75
SimpleBench41.4%—
NYT Connections (extended)—85.1%
CritPt—26%
Chess Puzzles—38%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
Mystery Game Puzzles—25%
LiveBench Data Analysis55%—
LMCA—53.9%
Bench to the Future 3—0.14
ForecastBench60.7—
LiveBench59%—

Math Muse Spark 1.3 leads

Claude 3.5 Sonnet: 19.2 (#288), Muse Spark 1.3: 73.1 (#21)

Math benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
OTIS Mock AIME 2024-20258.5%99.2%
LMArena Math13071494
FrontierMath (Tiers 1-3)—74.4%
FrontierMath Tier 4—46.3%
ProofBench—58%
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Muse Spark 1.3 leads

Claude 3.5 Sonnet: 28.6 (#245), Muse Spark 1.3: 42.6 (#95)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Expert12651516
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Muse Spark 1.3 leads

Claude 3.5 Sonnet: 26.5 (#120), Muse Spark 1.3: 43.7 (#22)

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Vision11251309
Video-MME60%—
GeoBench62%—
VPCT33%—
LMArena Document—1471

Multilingual Muse Spark 1.3 leads

Claude 3.5 Sonnet: 43.2 (#185), Muse Spark 1.3: 57.4 (#8)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Non-English12831481
LMArena Chinese12721529
LMArena French13051524
LMArena German12971515
LMArena Japanese12341474
LMArena Korean12001501
LMArena Russian13061490
LMArena Spanish12901490

Instruction Following Muse Spark 1.3 leads

Claude 3.5 Sonnet: 68.8 (#182), Muse Spark 1.3: 77.5 (#22)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Instruction Following12971477
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Muse Spark 1.3 leads

Claude 3.5 Sonnet: 39.9 (#167), Muse Spark 1.3: 45.6 (#32)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Longer Query13111488

Writing & Preference Muse Spark 1.3 leads

Claude 3.5 Sonnet: 52.9 (#164), Muse Spark 1.3: 73.6 (#9)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.3
LMArena Text12981490
LMArena Creative Writing12921455
EQ-Bench Creative Writing14511906
LMArena Multi-Turn13261482
Short-Story Creative Writing80.3%—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Muse Spark 1.3?

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Muse Spark 1.3 better for coding?

Muse Spark 1.3 scores higher on coding benchmarks: 56.6 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Muse Spark 1.3 share?

22 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Muse Spark 1.3 has 37.

Related comparisons

Go deeper