Model comparison

Muse Spark 1.3 vs Qwen Max

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 34.7 on the Noometry Index.

Last verified . 18 shared benchmarks.

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Muse Spark 1.3 scores higher in 8 categories and Qwen Max in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Muse Spark 1.3 leads 73.1 to 22.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 99.2% for Muse Spark 1.3 and 16.1% for Qwen Max.
  • Muse Spark 1.3 is cheaper at $1.25 / $4.25 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Muse Spark 1.3 accepts more context: 1.05M tokens versus 33K.

Side by side

Muse Spark 1.3 and Qwen Max specifications
Muse Spark 1.3Qwen Max
ProviderMetaAlibaba (Qwen)
Noometry Index54.834.7
Released2026-09-022024-04-03
WeightsProprietaryProprietary
Context window1.05M33K
Max output131K8K
Input $ / M tokens$1.25$1.60
Output $ / M tokens$4.25$6.40
Results tracked3723

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.3 leads

Muse Spark 1.3: 56.6 (#21), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Coding15141288
Aider Polyglot—21.8%
CursorBench41.6%—
LMArena WebDev1657—
SciCode59.7%—

Agentic & Tool Use Not comparable

Muse Spark 1.3: 38.6 (#30), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkMuse Spark 1.3Qwen Max
APEX-Agents57.8%—
GDP.pdf27.6%—

Reasoning Muse Spark 1.3 leads

Muse Spark 1.3: 54.0 (#27), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Hard Prompts15031269
NYT Connections (extended)85.1%—
CritPt26%—
Chess Puzzles38%—
Mystery Game Puzzles25%—
DTBench96.5%—
LMCA53.9%—
Bench to the Future 30.14—
Epoch Capabilities Index156.75—

Math Muse Spark 1.3 leads

Muse Spark 1.3: 73.1 (#21), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkMuse Spark 1.3Qwen Max
OTIS Mock AIME 2024-202599.2%16.1%
LMArena Math14941275
FrontierMath (Tiers 1-3)74.4%—
FrontierMath Tier 446.3%—
ProofBench58%—
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Muse Spark 1.3 leads

Muse Spark 1.3: 42.6 (#95), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Expert15161248
GPQA Diamond—56.1%

Multimodal Not comparable

Muse Spark 1.3: 43.7 (#22), Qwen Max: —

Multimodal benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Vision1309—
LMArena Document1471—

Multilingual Muse Spark 1.3 leads

Muse Spark 1.3: 57.4 (#8), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Non-English14811263
LMArena Chinese15291254
LMArena French15241330
LMArena German15151254
LMArena Japanese14741205
LMArena Korean15011142
LMArena Russian14901274
LMArena Spanish14901290

Instruction Following Muse Spark 1.3 leads

Muse Spark 1.3: 77.5 (#22), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Instruction Following14771262

Long Context Muse Spark 1.3 leads

Muse Spark 1.3: 45.6 (#32), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Longer Query14881288
Fiction.LiveBench—66.7%

Writing & Preference Muse Spark 1.3 leads

Muse Spark 1.3: 73.6 (#9), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkMuse Spark 1.3Qwen Max
LMArena Text14901282
LMArena Creative Writing14551248
LMArena Multi-Turn14821277
EQ-Bench Creative Writing1906—

Frequently asked questions

Is Muse Spark 1.3 better than Qwen Max?

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 34.7 on the Noometry Index.

Which is cheaper, Muse Spark 1.3 or Qwen Max?

Muse Spark 1.3 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Muse Spark 1.3 or Qwen Max better for coding?

Muse Spark 1.3 scores higher on coding benchmarks: 56.6 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.3 does, with 1.05M tokens against 33K.

How many benchmarks do Muse Spark 1.3 and Qwen Max share?

18 benchmarks have published results for both models. Muse Spark 1.3 has 37 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper