Model comparison

GLM-5.2 vs Muse Spark 1.2

GLM-5.2 and Muse Spark 1.2 score almost the same on the Noometry Index (51.1 vs 50.3), so choose on price, context window or the category you care about most.

Last verified . 28 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Summary

  • They share 28 benchmarks with published results for both. GLM-5.2 scores higher in 6 categories and Muse Spark 1.2 in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5.2 leads 55.7 to 46.4.
  • The biggest single-benchmark swing is SimpleQA Verified: 34.2% for GLM-5.2 and 60.3% for Muse Spark 1.2.
  • Muse Spark 1.2 is cheaper at $1.25 / $4.25 per million input/output tokens, against $1.40 / $4.40 for GLM-5.2.
  • Muse Spark 1.2 accepts more context: 1.05M tokens versus 1M.
  • GLM-5.2 has downloadable open weights; the other is API-only.

Side by side

GLM-5.2 and Muse Spark 1.2 specifications
GLM-5.2Muse Spark 1.2
ProviderZ.ai (Zhipu)Meta
Noometry Index51.150.3
Released2026-06-132026-08-05
WeightsOpenProprietary
Context window1M1.05M
Max output131K131K
Input $ / M tokens$1.40$1.25
Output $ / M tokens$4.40$4.25
Results tracked5131

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.2 leads

GLM-5.2: 51.3 (#41), Muse Spark 1.2: 49.2 (#51)

Coding benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
DeepSWE43.8%54.9%
LMArena WebDev16031533
SciCode50.5%56.4%
WeirdML70.1%60.3%
LMArena Coding14851495
SWE-bench Verified78.7%—
FrontierCode24.5%—
FrontierSWE—12%
ALE-Bench1,047—

Agentic & Tool Use GLM-5.2 leads

GLM-5.2: 32.4 (#63), Muse Spark 1.2: 29.4 (#87)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
APEX-Agents45.2%36.4%
τ²-bench Banking37.1%—
PostTrainBench31.7%—
GBAEval0%—
GDP.pdf—16%
Vending-Bench 28,314—

Reasoning Muse Spark 1.2 leads

GLM-5.2: 42.3 (#52), Muse Spark 1.2: 51.3 (#34)

Reasoning benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
SimpleBench58.8%74.5%
NYT Connections (extended)74.3%79.2%
CritPt20.9%17.7%
LMArena Hard Prompts14801486
DTBench93.6%94.7%
LMCA45.8%48.4%
Epoch Capabilities Index151.78154.87
ARC-AGI-222.8%—
Kagi LLM Benchmark62.6%—
ARC-AGI-177%—
Chess Puzzles21%—
EBR-Bench9.5%—
Mystery Game Puzzles19%—
Surface Evolver Bench55.6%—

Math GLM-5.2 leads

GLM-5.2: 55.7 (#43), Muse Spark 1.2: 46.4 (#70)

Math benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
ProofBench35%43%
LMArena Math14821471
FrontierMath (Tiers 1-3)59.2%—
FrontierMath Tier 429.3%—
MathArena Final-Answer Competitions67.6%—
OTIS Mock AIME 2024-202586.4%—

Knowledge GLM-5.2 leads

GLM-5.2: 57.1 (#40), Muse Spark 1.2: 54.1 (#53)

Knowledge benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
SimpleQA Verified34.2%60.3%
LMArena Expert14861480
GPQA Diamond91.9%—

Multimodal Not comparable

GLM-5.2: —, Muse Spark 1.2: 43.4 (#25)

Multimodal benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
LMArena Vision—1305

Multilingual Muse Spark 1.2 leads

GLM-5.2: 55.8 (#26), Muse Spark 1.2: 57.1 (#11)

Multilingual benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
LMArena Non-English14591478
LMArena Chinese15191511
LMArena French14791513
LMArena Russian14661487
LMArena Spanish14771498
LMArena German1468—
LMArena Japanese1451—
LMArena Korean1445—

Instruction Following Too close to call

GLM-5.2: 76.9 (#34), Muse Spark 1.2: 76.7 (#36)

Instruction Following benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
LMArena Instruction Following14651461

Long Context Too close to call

GLM-5.2: 45.3 (#43), Muse Spark 1.2: 45.2 (#48)

Long Context benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
LMArena Longer Query14791475

Writing & Preference Muse Spark 1.2 leads

GLM-5.2: 70.4 (#21), Muse Spark 1.2: 72.3 (#14)

Writing & Preference benchmarks
BenchmarkGLM-5.2Muse Spark 1.2
LMArena Text14701482
LMArena Creative Writing14621449
EQ-Bench Creative Writing17571840
LMArena Multi-Turn14691494
EQ-Bench 41222—

Frequently asked questions

Is GLM-5.2 better than Muse Spark 1.2?

GLM-5.2 and Muse Spark 1.2 score almost the same on the Noometry Index (51.1 vs 50.3), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.2 or Muse Spark 1.2?

Muse Spark 1.2 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; GLM-5.2 lists at $1.40 and $4.40.

Is GLM-5.2 or Muse Spark 1.2 better for coding?

GLM-5.2 scores higher on coding benchmarks: 51.3 versus 49.2 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.2 does, with 1.05M tokens against 1M.

How many benchmarks do GLM-5.2 and Muse Spark 1.2 share?

28 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Muse Spark 1.2 has 31.

Related comparisons

Go deeper