Model comparison

ERNIE 5.1 vs Muse Spark 1.3

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 43.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Summary

  • They share 18 benchmarks with published results for both. ERNIE 5.1 scores higher in 0 categories and Muse Spark 1.3 in 8 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Muse Spark 1.3 leads 73.1 to 40.3.
  • The biggest single-benchmark swing is NYT Connections (extended): 23.4% for ERNIE 5.1 and 85.1% for Muse Spark 1.3.

Side by side

ERNIE 5.1 and Muse Spark 1.3 specifications
ERNIE 5.1Muse Spark 1.3
ProviderBaiduMeta
Noometry Index43.854.8
Released—2026-09-02
WeightsProprietaryProprietary
Context window—1.05M
Max output—131K
Input $ / M tokens—$1.25
Output $ / M tokens—$4.25
Results tracked1937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.3 leads

ERNIE 5.1: 44.0 (#76), Muse Spark 1.3: 56.6 (#21)

Coding benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Coding14881514
CursorBench—41.6%
LMArena WebDev—1657
SciCode—59.7%

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Muse Spark 1.3: 38.6 (#30)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
APEX-Agents—57.8%
GDP.pdf—27.6%
LMArena Search1227—

Reasoning Muse Spark 1.3 leads

ERNIE 5.1: 21.9 (#211), Muse Spark 1.3: 54.0 (#27)

Reasoning benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
NYT Connections (extended)23.4%85.1%
LMArena Hard Prompts14811503
CritPt—26%
Chess Puzzles—38%
Mystery Game Puzzles—25%
DTBench—96.5%
LMCA—53.9%
Bench to the Future 3—0.14
Epoch Capabilities Index—156.75

Math Muse Spark 1.3 leads

ERNIE 5.1: 40.3 (#92), Muse Spark 1.3: 73.1 (#21)

Math benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Math14811494
FrontierMath (Tiers 1-3)—74.4%
FrontierMath Tier 4—46.3%
OTIS Mock AIME 2024-2025—99.2%
ProofBench—58%

Knowledge Too close to call

ERNIE 5.1: 41.9 (#102), Muse Spark 1.3: 42.6 (#95)

Knowledge benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Expert14921516

Multimodal Not comparable

ERNIE 5.1: —, Muse Spark 1.3: 43.7 (#22)

Multimodal benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Vision—1309
LMArena Document—1471

Multilingual Muse Spark 1.3 leads

ERNIE 5.1: 55.5 (#29), Muse Spark 1.3: 57.4 (#8)

Multilingual benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Non-English14541481
LMArena Chinese15081529
LMArena French14881524
LMArena German14701515
LMArena Japanese14221474
LMArena Korean14271501
LMArena Russian14591490
LMArena Spanish14731490

Instruction Following Too close to call

ERNIE 5.1: 76.7 (#37), Muse Spark 1.3: 77.5 (#22)

Instruction Following benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Instruction Following14601477

Long Context Too close to call

ERNIE 5.1: 44.7 (#59), Muse Spark 1.3: 45.6 (#32)

Long Context benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Longer Query14621488

Writing & Preference Muse Spark 1.3 leads

ERNIE 5.1: 65.1 (#52), Muse Spark 1.3: 73.6 (#9)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Muse Spark 1.3
LMArena Text14681490
LMArena Creative Writing14411455
LMArena Multi-Turn14711482
EQ-Bench Creative Writing—1906

Frequently asked questions

Is ERNIE 5.1 better than Muse Spark 1.3?

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 43.8 on the Noometry Index.

Is ERNIE 5.1 or Muse Spark 1.3 better for coding?

Muse Spark 1.3 scores higher on coding benchmarks: 56.6 versus 44.0 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Muse Spark 1.3 share?

18 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Muse Spark 1.3 has 37.

Related comparisons

Go deeper