Model comparison

Longcat Flash Chat vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 42.1 on the Noometry Index.

Last verified . 16 shared benchmarks.

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Longcat Flash Chat scores higher in 0 categories and Muse Spark in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 40.6.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Longcat Flash Chat and Muse Spark specifications
Longcat Flash ChatMuse Spark
ProviderMeituanMeta
Noometry Index42.150.6
Released—2026-04-08
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

Longcat Flash Chat: 43.5 (#87), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Coding14711481
SciCode—51.5%

Reasoning Muse Spark leads

Longcat Flash Chat: 19.0 (#272), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Hard Prompts14401474
Kagi LLM Benchmark43.9%—
NYT Connections (extended)17.7%—
CritPt—11.3%
Epoch Capabilities Index—152.04

Math Muse Spark leads

Longcat Flash Chat: 39.4 (#107), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Math14421455
OTIS Mock AIME 2024-2025—88.9%
ProofBench—17%
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

Longcat Flash Chat: 40.6 (#116), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Expert14541457
GPQA Diamond—89.8%
Humanity's Last Exam—40.6%

Multimodal Not comparable

Longcat Flash Chat: —, Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Vision—1306
LMArena Document—1444

Multilingual Muse Spark leads

Longcat Flash Chat: 51.9 (#101), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Non-English14041464
LMArena Chinese14651509
LMArena French14561497
LMArena German14081497
LMArena Korean13711459
LMArena Russian13951466
LMArena Spanish14451472
LMArena Japanese1373—

Instruction Following Muse Spark leads

Longcat Flash Chat: 74.4 (#96), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Instruction Following14111442

Long Context Too close to call

Longcat Flash Chat: 43.5 (#93), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Longer Query14251451

Writing & Preference Muse Spark leads

Longcat Flash Chat: 61.0 (#91), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkLongcat Flash ChatMuse Spark
LMArena Text14271474
LMArena Creative Writing13881459
LMArena Multi-Turn14181477

Frequently asked questions

Is Longcat Flash Chat better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 42.1 on the Noometry Index.

Is Longcat Flash Chat or Muse Spark better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 43.5 in the Noometry coding category.

How many benchmarks do Longcat Flash Chat and Muse Spark share?

16 benchmarks have published results for both models. Longcat Flash Chat has 19 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper