Model comparison

Muse Spark 1.3 vs Phi-4

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 31.2 on the Noometry Index. Phi-4 costs 23× less per token, which makes it the better buy when Muse Spark 1.3's lead doesn't matter for your workload.

Last verified . 20 shared benchmarks.

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Muse Spark 1.3 scores higher in 9 categories and Phi-4 in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Muse Spark 1.3 leads 73.1 to 20.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 99.2% for Muse Spark 1.3 and 13.8% for Phi-4.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $1.25 / $4.25 for Muse Spark 1.3.
  • Muse Spark 1.3 accepts more context: 1.05M tokens versus 128K.
  • Phi-4 has downloadable open weights; the other is API-only.

Side by side

Muse Spark 1.3 and Phi-4 specifications
Muse Spark 1.3Phi-4
ProviderMetaMicrosoft
Noometry Index54.831.2
Released2026-09-022024-12-11
WeightsProprietaryOpen
Context window1.05M128K
Max output131K4K
Input $ / M tokens$1.25$0.07
Output $ / M tokens$4.25$0.14
Results tracked3737

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.3 leads

Muse Spark 1.3: 56.6 (#21), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkMuse Spark 1.3Phi-4
LMArena Coding15141231
CursorBench41.6%—
LMArena WebDev1657—
SciCode59.7%—
BigCodeBench Instruct—45.5%
LiveBench Coding—30.7%
BigCodeBench Complete—55.4%

Agentic & Tool Use Muse Spark 1.3 leads

Muse Spark 1.3: 38.6 (#30), Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkMuse Spark 1.3Phi-4
APEX-Agents57.8%—
Berkeley Function Calling Leaderboard—28.8%
BALROG—11.6%
GDP.pdf27.6%—

Reasoning Muse Spark 1.3 leads

Muse Spark 1.3: 54.0 (#27), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkMuse Spark 1.3Phi-4
Chess Puzzles38%1%
LMArena Hard Prompts15031220
Epoch Capabilities Index156.75130.42
NYT Connections (extended)85.1%—
CritPt26%—
LiveBench Reasoning—47.8%
Mystery Game Puzzles25%—
DTBench96.5%—
LiveBench Data Analysis—45.2%
LMCA53.9%—
Bench to the Future 30.14—
LiveBench—41.6%

Math Muse Spark 1.3 leads

Muse Spark 1.3: 73.1 (#21), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkMuse Spark 1.3Phi-4
OTIS Mock AIME 2024-202599.2%13.8%
LMArena Math14941246
FrontierMath (Tiers 1-3)74.4%—
FrontierMath Tier 446.3%—
ProofBench58%—
LiveBench Math—42%
MATH Level 5—64.9%

Knowledge Muse Spark 1.3 leads

Muse Spark 1.3: 42.6 (#95), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkMuse Spark 1.3Phi-4
LMArena Expert15161203
GPQA Diamond—56.1%
Confabulations—29.4%
Vectara Hallucination Rate—3.7%
MMLU—84.8%

Multimodal Not comparable

Muse Spark 1.3: 43.7 (#22), Phi-4: —

Multimodal benchmarks
BenchmarkMuse Spark 1.3Phi-4
LMArena Vision1309—
LMArena Document1471—

Multilingual Muse Spark 1.3 leads

Muse Spark 1.3: 57.4 (#8), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkMuse Spark 1.3Phi-4
LMArena Non-English14811197
LMArena Chinese15291212
LMArena French15241224
LMArena German15151222
LMArena Japanese14741158
LMArena Korean15011151
LMArena Russian14901209
LMArena Spanish14901234

Instruction Following Muse Spark 1.3 leads

Muse Spark 1.3: 77.5 (#22), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkMuse Spark 1.3Phi-4
LMArena Instruction Following14771201
LiveBench Instruction Following—58.4%

Long Context Muse Spark 1.3 leads

Muse Spark 1.3: 45.6 (#32), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkMuse Spark 1.3Phi-4
LMArena Longer Query14881217

Writing & Preference Muse Spark 1.3 leads

Muse Spark 1.3: 73.6 (#9), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkMuse Spark 1.3Phi-4
LMArena Text14901217
LMArena Creative Writing14551182
LMArena Multi-Turn14821206
Short-Story Creative Writing—62.6%
EQ-Bench Creative Writing1906—
LiveBench Language—25.6%

Frequently asked questions

Is Muse Spark 1.3 better than Phi-4?

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 31.2 on the Noometry Index. Phi-4 costs 23× less per token, which makes it the better buy when Muse Spark 1.3's lead doesn't matter for your workload.

Which is cheaper, Muse Spark 1.3 or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Muse Spark 1.3 lists at $1.25 and $4.25.

Is Muse Spark 1.3 or Phi-4 better for coding?

Muse Spark 1.3 scores higher on coding benchmarks: 56.6 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.3 does, with 1.05M tokens against 128K.

How many benchmarks do Muse Spark 1.3 and Phi-4 share?

20 benchmarks have published results for both models. Muse Spark 1.3 has 37 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper