Model comparison

GLM-5.2 vs Muse Spark

GLM-5.2 and Muse Spark score almost the same on the Noometry Index (51.1 vs 50.6), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 22 benchmarks with published results for both. GLM-5.2 scores higher in 6 categories and Muse Spark in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 57.1.
  • The biggest single-benchmark swing is ProofBench: 35% for GLM-5.2 and 17% for Muse Spark.
  • GLM-5.2 has downloadable open weights; the other is API-only.

Side by side

GLM-5.2 and Muse Spark specifications
GLM-5.2Muse Spark
ProviderZ.ai (Zhipu)Meta
Noometry Index51.150.6
Released2026-06-132026-04-08
WeightsOpenProprietary
Context window1M—
Max output131K—
Input $ / M tokens$1.40—
Output $ / M tokens$4.40—
Results tracked5127

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.2 leads

GLM-5.2: 51.3 (#41), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGLM-5.2Muse Spark
SciCode50.5%51.5%
LMArena Coding14851481
SWE-bench Verified78.7%—
DeepSWE43.8%—
FrontierCode24.5%—
LMArena WebDev1603—
WeirdML70.1%—
ALE-Bench1,047—

Agentic & Tool Use Not comparable

GLM-5.2: 32.4 (#63), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2Muse Spark
APEX-Agents45.2%—
τ²-bench Banking37.1%—
PostTrainBench31.7%—
GBAEval0%—
Vending-Bench 28,314—

Reasoning GLM-5.2 leads

GLM-5.2: 42.3 (#52), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGLM-5.2Muse Spark
CritPt20.9%11.3%
LMArena Hard Prompts14801474
Epoch Capabilities Index151.78152.04
ARC-AGI-222.8%—
SimpleBench58.8%—
Kagi LLM Benchmark62.6%—
NYT Connections (extended)74.3%—
ARC-AGI-177%—
Chess Puzzles21%—
EBR-Bench9.5%—
Mystery Game Puzzles19%—
DTBench93.6%—
LMCA45.8%—
Surface Evolver Bench55.6%—

Math GLM-5.2 leads

GLM-5.2: 55.7 (#43), Muse Spark: 47.8 (#66)

Knowledge Muse Spark leads

GLM-5.2: 57.1 (#40), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGLM-5.2Muse Spark
GPQA Diamond91.9%89.8%
LMArena Expert14861457
Humanity's Last Exam—40.6%
SimpleQA Verified34.2%—

Multimodal Not comparable

GLM-5.2: —, Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGLM-5.2Muse Spark
LMArena Vision—1306
LMArena Document—1444

Multilingual Too close to call

GLM-5.2: 55.8 (#26), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGLM-5.2Muse Spark
LMArena Non-English14591464
LMArena Chinese15191509
LMArena French14791497
LMArena German14681497
LMArena Korean14451459
LMArena Russian14661466
LMArena Spanish14771472
LMArena Japanese1451—

Instruction Following GLM-5.2 leads

GLM-5.2: 76.9 (#34), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGLM-5.2Muse Spark
LMArena Instruction Following14651442

Long Context Too close to call

GLM-5.2: 45.3 (#43), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGLM-5.2Muse Spark
LMArena Longer Query14791451

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGLM-5.2Muse Spark
LMArena Text14701474
LMArena Creative Writing14621459
LMArena Multi-Turn14691477
EQ-Bench Creative Writing1757—
EQ-Bench 41222—

Frequently asked questions

Is GLM-5.2 better than Muse Spark?

GLM-5.2 and Muse Spark score almost the same on the Noometry Index (51.1 vs 50.6), so choose on price, context window or the category you care about most.

Is GLM-5.2 or Muse Spark better for coding?

GLM-5.2 scores higher on coding benchmarks: 51.3 versus 46.2 in the Noometry coding category.

How many benchmarks do GLM-5.2 and Muse Spark share?

22 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper