Model comparison

DeepSeek V4 Flash vs Muse Spark

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 50.6 on the Noometry Index.

Last verified . 22 shared benchmarks.

DeepSeek V4 Flash DeepSeek

53.6

Rank #35 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 22 benchmarks with published results for both. DeepSeek V4 Flash scores higher in 3 categories and Muse Spark in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Flash leads 53.7 to 35.9.
  • The biggest single-benchmark swing is ProofBench: 56% for DeepSeek V4 Flash and 17% for Muse Spark.
  • DeepSeek V4 Flash has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4 Flash and Muse Spark specifications
DeepSeek V4 FlashMuse Spark
ProviderDeepSeekMeta
Noometry Index53.650.6
Released2026-04-242026-04-08
WeightsOpenProprietary
Context window1M—
Max output393K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked4127

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Flash leads

DeepSeek V4 Flash: 47.9 (#59), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
SciCode49.9%51.5%
LMArena Coding14571481
FrontierCode18.8%—
LMArena WebDev1582—
WeirdML63%—
ALE-Bench1,306—

Reasoning DeepSeek V4 Flash leads

DeepSeek V4 Flash: 53.7 (#30), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
CritPt16.6%11.3%
LMArena Hard Prompts14441474
Epoch Capabilities Index154.49152.04
ARC-AGI-261.4%—
SimpleBench61.1%—
Kagi LLM Benchmark52.2%—
NYT Connections (extended)89.6%—
ARC-AGI-189%—
Chess Puzzles33%—
Mystery Game Puzzles34%—
DTBench90.9%—
LMCA41.7%—

Math DeepSeek V4 Flash leads

DeepSeek V4 Flash: 60.3 (#37), Muse Spark: 47.8 (#66)

Knowledge Muse Spark leads

DeepSeek V4 Flash: 55.4 (#48), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
GPQA Diamond91%89.8%
LMArena Expert14411457
Humanity's Last Exam—40.6%
SimpleQA Verified33.6%—

Multimodal Not comparable

DeepSeek V4 Flash: —, Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
LMArena Vision—1306
LMArena Document—1444

Multilingual Muse Spark leads

DeepSeek V4 Flash: 53.0 (#72), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
LMArena Non-English14201464
LMArena Chinese14681509
LMArena French14391497
LMArena German14181497
LMArena Korean13841459
LMArena Russian14281466
LMArena Spanish14361472
LMArena Japanese1406—

Instruction Following Muse Spark leads

DeepSeek V4 Flash: 74.9 (#81), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
LMArena Instruction Following14211442

Long Context Too close to call

DeepSeek V4 Flash: 43.8 (#85), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
LMArena Longer Query14341451

Writing & Preference Muse Spark leads

DeepSeek V4 Flash: 63.8 (#61), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkDeepSeek V4 FlashMuse Spark
LMArena Text14321474
LMArena Creative Writing14031459
LMArena Multi-Turn14491477
EQ-Bench Creative Writing1559—

Frequently asked questions

Is DeepSeek V4 Flash better than Muse Spark?

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 50.6 on the Noometry Index.

Is DeepSeek V4 Flash or Muse Spark better for coding?

DeepSeek V4 Flash scores higher on coding benchmarks: 47.9 versus 46.2 in the Noometry coding category.

How many benchmarks do DeepSeek V4 Flash and Muse Spark share?

22 benchmarks have published results for both models. DeepSeek V4 Flash has 41 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper