Model comparison

GLM-5.3 vs Nemotron 3 Ultra

GLM-5.3 is the stronger model overall, scoring 54.8 to 42.5 on the Noometry Index. Nemotron 3 Ultra costs 2.3× less per token, which makes it the better buy when GLM-5.3's lead doesn't matter for your workload.

Last verified . 30 shared benchmarks.

GLM-5.3 Z.ai (Zhipu)

54.8

Rank #26 Confirmed

Nemotron 3 Ultra NVIDIA

42.5

Rank #113 Confirmed

Summary

  • They share 30 benchmarks with published results for both. GLM-5.3 scores higher in 9 categories and Nemotron 3 Ultra in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5.3 leads 62.3 to 35.3.
  • The biggest single-benchmark swing is ProofBench: 49% for GLM-5.3 and 2% for Nemotron 3 Ultra.
  • Nemotron 3 Ultra is cheaper at $0.50 / $2.20 per million input/output tokens, against $1.40 / $4.40 for GLM-5.3.
  • GLM-5.3 accepts more context: 1M tokens versus 262K.

Side by side

GLM-5.3 and Nemotron 3 Ultra specifications
GLM-5.3Nemotron 3 Ultra
ProviderZ.ai (Zhipu)NVIDIA
Noometry Index54.842.5
Released2026-08-142026-06-04
WeightsOpenOpen
Context window1M262K
Max output131K128K
Input $ / M tokens$1.40$0.50
Output $ / M tokens$4.40$2.20
Results tracked4230

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3 leads

GLM-5.3: 59.5 (#14), Nemotron 3 Ultra: 38.1 (#182)

Coding benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
FrontierCode40.1%13.6%
SciCode59%40.3%
WeirdML75.4%43.5%
LMArena Coding14961468
DeepSWE69%—
CursorBench42.6%—
LMArena WebDev1622—
FrontierSWE30.2%—
ALE-Bench1,317—

Agentic & Tool Use GLM-5.3 leads

GLM-5.3: 36.4 (#38), Nemotron 3 Ultra: 23.4 (#126)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
APEX-Agents56.6%22.7%
Vending-Bench 28,164—

Reasoning GLM-5.3 leads

GLM-5.3: 46.1 (#46), Nemotron 3 Ultra: 31.0 (#82)

Reasoning benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
CritPt19.1%3.1%
Chess Puzzles21%12%
LMArena Hard Prompts14891452
Mystery Game Puzzles33%20%
DTBench87.7%90.1%
LMCA55.5%36.9%
Epoch Capabilities Index155.61146.17
NYT Connections (extended)74.2%—
Bench to the Future 30.15—

Math GLM-5.3 leads

GLM-5.3: 62.3 (#33), Nemotron 3 Ultra: 35.3 (#187)

Math benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
OTIS Mock AIME 2024-202591.1%86.7%
ProofBench49%2%
LMArena Math14891457
FrontierMath (Tiers 1-3)68.8%—
FrontierMath Tier 429.3%—

Knowledge GLM-5.3 leads

GLM-5.3: 58.3 (#37), Nemotron 3 Ultra: 52.5 (#61)

Knowledge benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
GPQA Diamond90.9%85.4%
LMArena Expert15161472
SimpleQA Verified41%—

Multilingual GLM-5.3 leads

GLM-5.3: 55.7 (#28), Nemotron 3 Ultra: 53.5 (#64)

Multilingual benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
LMArena Non-English14571427
LMArena Chinese15281497
LMArena French14991472
LMArena German14991471
LMArena Korean14721386
LMArena Russian14631417
LMArena Spanish14601454
LMArena Japanese1453—

Instruction Following GLM-5.3 leads

GLM-5.3: 77.5 (#23), Nemotron 3 Ultra: 74.7 (#88)

Instruction Following benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
LMArena Instruction Following14771418

Long Context GLM-5.3 leads

GLM-5.3: 45.4 (#41), Nemotron 3 Ultra: 43.9 (#81)

Long Context benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
LMArena Longer Query14821437

Writing & Preference GLM-5.3 leads

GLM-5.3: 75.7 (#6), Nemotron 3 Ultra: 66.3 (#36)

Writing & Preference benchmarks
BenchmarkGLM-5.3Nemotron 3 Ultra
LMArena Text14711445
LMArena Creative Writing14571398
EQ-Bench Creative Writing20751692
LMArena Multi-Turn14721411

Frequently asked questions

Is GLM-5.3 better than Nemotron 3 Ultra?

GLM-5.3 is the stronger model overall, scoring 54.8 to 42.5 on the Noometry Index. Nemotron 3 Ultra costs 2.3× less per token, which makes it the better buy when GLM-5.3's lead doesn't matter for your workload.

Which is cheaper, GLM-5.3 or Nemotron 3 Ultra?

Nemotron 3 Ultra is cheaper. It lists at $0.50 per million input tokens and $2.20 per million output tokens; GLM-5.3 lists at $1.40 and $4.40.

Is GLM-5.3 or Nemotron 3 Ultra better for coding?

GLM-5.3 scores higher on coding benchmarks: 59.5 versus 38.1 in the Noometry coding category.

Which has the bigger context window?

GLM-5.3 does, with 1M tokens against 262K.

How many benchmarks do GLM-5.3 and Nemotron 3 Ultra share?

30 benchmarks have published results for both models. GLM-5.3 has 42 scored results on Noometry and Nemotron 3 Ultra has 30.

Related comparisons

Go deeper