Model comparison

Gemini 2.5 Pro vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 45.0 on the Noometry Index.

Last verified . 26 shared benchmarks.

Gemini 2.5 Pro Google

45.0

Rank #75 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Gemini 2.5 Pro scores higher in 2 categories and Muse Spark in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Gemini 2.5 Pro leads 59.8 to 44.4.
  • The biggest single-benchmark swing is Humanity's Last Exam: 21.6% for Gemini 2.5 Pro and 40.6% for Muse Spark.

Side by side

Gemini 2.5 Pro and Muse Spark specifications
Gemini 2.5 ProMuse Spark
ProviderGoogleMeta
Noometry Index45.050.6
Released2025-03-252026-04-08
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$1.25—
Output $ / M tokens$10—
Results tracked7827

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

Gemini 2.5 Pro: 42.4 (#101), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGemini 2.5 ProMuse Spark
SciCode42.8%51.5%
LMArena Coding14521481
SWE-bench Verified57.6%—
SWE-bench Verified (bash only)53.6%—
Aider Polyglot83.1%—
LMArena WebDev1227—
GSO3.9%—
WeirdML54%—
LiveBench Coding85.9%—
CadEval64%—
ALE-Bench785.52—
AlgoTune1.51—

Agentic & Tool Use Not comparable

Gemini 2.5 Pro: 29.2 (#88), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 ProMuse Spark
Terminal-Bench32.6%—
GDPval23.3%—
Remote Labor Index0.8%—
TheAgentCompany30.3%—
τ²-bench Banking13.7%—
DeepResearch Bench42.8%—
BALROG43.3%—
LMArena Search1142—
METR Time Horizons55.4%—
Vending-Bench 2573.64—

Reasoning Muse Spark leads

Gemini 2.5 Pro: 28.8 (#99), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGemini 2.5 ProMuse Spark
CritPt2%11.3%
LMArena Hard Prompts14551474
Epoch Capabilities Index145.32152.04
ARC-AGI-24.9%—
SimpleBench62.4%—
Kagi LLM Benchmark70.3%—
ARC-AGI-141%—
Chess Puzzles20%—
EnigmaEval5.6%—
LiveBench Reasoning89.8%—
DTBench82.4%—
LiveBench Data Analysis79.9%—
LMCA34.8%—
ForecastBench61.3—
LiveBench82.3%—

Math Muse Spark leads

Gemini 2.5 Pro: 32.5 (#213), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGemini 2.5 ProMuse Spark
OTIS Mock AIME 2024-202584.7%88.9%
LMArena Math14501455
FrontierMath (Feb 2025 set)14.1%39%
FrontierMath Tier 4 (v1)4.2%14.6%
FrontierMath (Tiers 1-3)24.6%—
FrontierMath Tier 40%—
ProofBench—17%
Omni-MATH41.6%—
LiveBench Math90.2%—
MATH Level 595.9%—

Knowledge Muse Spark leads

Gemini 2.5 Pro: 56.0 (#46), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGemini 2.5 ProMuse Spark
GPQA Diamond85.3%89.8%
Humanity's Last Exam21.6%40.6%
LMArena Expert14521457
MMLU-Pro86.3%—
Confabulations10.6%—
Vectara Hallucination Rate7%—
GPQA (HELM)74.9%—

Multimodal Gemini 2.5 Pro leads

Gemini 2.5 Pro: 45.2 (#18), Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGemini 2.5 ProMuse Spark
LMArena Vision12631306
LMArena Document14211444
GeoBench86%—
VPCT48%—
SpatialViz-Bench44.7%—

Multilingual Too close to call

Gemini 2.5 Pro: 55.3 (#31), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGemini 2.5 ProMuse Spark
LMArena Non-English14511464
LMArena Chinese15071509
LMArena French14721497
LMArena German14871497
LMArena Korean14341459
LMArena Russian14611466
LMArena Spanish14731472
LMArena Japanese1461—

Instruction Following Too close to call

Gemini 2.5 Pro: 75.0 (#75), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGemini 2.5 ProMuse Spark
LMArena Instruction Following14371442
LiveBench Instruction Following80.6%—
IFEval84%—

Long Context Gemini 2.5 Pro leads

Gemini 2.5 Pro: 59.8 (#5), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGemini 2.5 ProMuse Spark
LMArena Longer Query14491451
Fiction.LiveBench91.7%—

Writing & Preference Muse Spark leads

Gemini 2.5 Pro: 63.7 (#62), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGemini 2.5 ProMuse Spark
LMArena Text14581474
LMArena Creative Writing14541459
LMArena Multi-Turn14531477
Short-Story Creative Writing83.8%—
EQ-Bench Creative Writing1421—
WildBench85.7%—
LiveBench Language67.8%—

Frequently asked questions

Is Gemini 2.5 Pro better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 45.0 on the Noometry Index.

Is Gemini 2.5 Pro or Muse Spark better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 42.4 in the Noometry coding category.

How many benchmarks do Gemini 2.5 Pro and Muse Spark share?

26 benchmarks have published results for both models. Gemini 2.5 Pro has 78 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper