Model comparison

Muse Spark vs Qwen3.6 Max Preview

Muse Spark and Qwen3.6 Max Preview score almost the same on the Noometry Index (50.6 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 19 shared benchmarks.

Muse Spark Meta

50.6

Rank #46 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Muse Spark scores higher in 4 categories and Qwen3.6 Max Preview in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 57.6.

Side by side

Muse Spark and Qwen3.6 Max Preview specifications
Muse SparkQwen3.6 Max Preview
ProviderMetaAlibaba (Qwen)
Noometry Index50.651.5
Released2026-04-082026-04-20
WeightsProprietaryProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.30
Output $ / M tokens—$7.80
Results tracked2729

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Muse Spark: 46.2 (#69), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
LMArena Coding14811471
SWE-bench Verified—76.7%
LMArena WebDev—1482
SciCode51.5%—

Agentic & Tool Use Not comparable

Muse Spark: —, Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Muse Spark: 35.9 (#67), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
LMArena Hard Prompts14741457
Epoch Capabilities Index152.04149.24
SimpleBench—63%
NYT Connections (extended)—74.1%
CritPt11.3%—
Chess Puzzles—20%
Mystery Game Puzzles—19%
DTBench—87.2%
LMCA—42.5%

Math Qwen3.6 Max Preview leads

Muse Spark: 47.8 (#66), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
OTIS Mock AIME 2024-202588.9%91.1%
LMArena Math14551465
FrontierMath (Feb 2025 set)39%23.1%
FrontierMath Tier 4 (v1)14.6%4.2%
ProofBench17%—

Knowledge Muse Spark leads

Muse Spark: 65.7 (#13), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
GPQA Diamond89.8%87.4%
LMArena Expert14571478
Humanity's Last Exam40.6%—
SimpleQA Verified—52%

Multimodal Not comparable

Muse Spark: 43.4 (#24), Qwen3.6 Max Preview: —

Multimodal benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
LMArena Vision1306—
LMArena Document1444—

Multilingual Muse Spark leads

Muse Spark: 56.1 (#24), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
LMArena Non-English14641437
LMArena Chinese15091487
LMArena French14971449
LMArena Russian14661445
LMArena Spanish14721454
LMArena German1497—
LMArena Korean1459—

Instruction Following Too close to call

Muse Spark: 75.9 (#51), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
LMArena Instruction Following14421438

Long Context Too close to call

Muse Spark: 44.4 (#69), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
LMArena Longer Query14511457

Writing & Preference Muse Spark leads

Muse Spark: 66.0 (#39), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkMuse SparkQwen3.6 Max Preview
LMArena Text14741447
LMArena Creative Writing14591435
LMArena Multi-Turn14771456

Frequently asked questions

Is Muse Spark better than Qwen3.6 Max Preview?

Muse Spark and Qwen3.6 Max Preview score almost the same on the Noometry Index (50.6 vs 51.5), so choose on price, context window or the category you care about most.

Is Muse Spark or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 46.2 in the Noometry coding category.

How many benchmarks do Muse Spark and Qwen3.6 Max Preview share?

19 benchmarks have published results for both models. Muse Spark has 27 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper