Model comparison

Claude Sonnet 4.6 vs GLM-5.3

GLM-5.3 is the stronger model overall, scoring 54.8 to 50.3 on the Noometry Index.

Last verified . 37 shared benchmarks.

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

GLM-5.3 Z.ai (Zhipu)

54.8

Rank #26 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Claude Sonnet 4.6 scores higher in 2 categories and GLM-5.3 in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GLM-5.3 leads 59.5 to 46.3.
  • The biggest single-benchmark swing is DeepSWE: 29.9% for Claude Sonnet 4.6 and 69% for GLM-5.3.
  • GLM-5.3 is cheaper at $1.40 / $4.40 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.6.
  • GLM-5.3 has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4.6 and GLM-5.3 specifications
Claude Sonnet 4.6GLM-5.3
ProviderAnthropicZ.ai (Zhipu)
Noometry Index50.354.8
Released2026-02-172026-08-14
WeightsProprietaryOpen
Context window1M1M
Max output128K131K
Input $ / M tokens$3$1.40
Output $ / M tokens$15$4.40
Results tracked5742

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3 leads

Claude Sonnet 4.6: 46.3 (#67), GLM-5.3: 59.5 (#14)

Coding benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
DeepSWE29.9%69%
FrontierCode24.3%40.1%
LMArena WebDev15221622
SciCode46.8%59%
WeirdML66.1%75.4%
LMArena Coding15041496
ALE-Bench1,3271,317
SWE-bench Verified75.2%—
CursorBench—42.6%
FrontierSWE—30.2%

Agentic & Tool Use Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 39.1 (#28), GLM-5.3: 36.4 (#38)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
APEX-Agents43%56.6%
Vending-Bench 27,2048,164
Terminal-Bench53.4%—
OSWorld 2.09.3%—
DeepResearch Bench54.9%—
OSWorld72.1%—
ExploitBench23.6%—
GBAEval48.8%—
GDP.pdf18%—
LMArena Search1221—

Reasoning Too close to call

Claude Sonnet 4.6: 46.1 (#45), GLM-5.3: 46.1 (#46)

Reasoning benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
NYT Connections (extended)80.9%74.2%
CritPt3.1%19.1%
Chess Puzzles13%21%
LMArena Hard Prompts14841489
Mystery Game Puzzles16%33%
DTBench89.9%87.7%
LMCA46.5%55.5%
Epoch Capabilities Index152.24155.61
ARC-AGI-260.4%—
ARC-AGI-186.5%—
Thematic Generalization76.3%—
Bench to the Future 3—0.15
ForecastBench62—

Math GLM-5.3 leads

Claude Sonnet 4.6: 52.9 (#49), GLM-5.3: 62.3 (#33)

Math benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
OTIS Mock AIME 2024-202585.8%91.1%
ProofBench45%49%
LMArena Math14621489
FrontierMath (Tiers 1-3)—68.8%
FrontierMath Tier 4—29.3%
FrontierMath (Feb 2025 set)32.4%—
FrontierMath Tier 4 (v1)8.3%—

Knowledge GLM-5.3 leads

Claude Sonnet 4.6: 51.7 (#65), GLM-5.3: 58.3 (#37)

Knowledge benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
GPQA Diamond87.4%90.9%
SimpleQA Verified35.5%41%
LMArena Expert15001516
Vectara Hallucination Rate10.6%—

Multimodal Not comparable

Claude Sonnet 4.6: 38.0 (#68), GLM-5.3: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
LMArena Vision1283—
Blueprint-Bench 26.7%—
LMArena Document1482—

Multilingual GLM-5.3 leads

Claude Sonnet 4.6: 54.4 (#41), GLM-5.3: 55.7 (#28)

Multilingual benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
LMArena Non-English14401457
LMArena Chinese14911528
LMArena French14651499
LMArena German14281499
LMArena Japanese14201453
LMArena Korean14111472
LMArena Russian14401463
LMArena Spanish14641460

Instruction Following Too close to call

Claude Sonnet 4.6: 77.4 (#25), GLM-5.3: 77.5 (#23)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
LMArena Instruction Following14751477

Long Context Too close to call

Claude Sonnet 4.6: 45.3 (#44), GLM-5.3: 45.4 (#41)

Long Context benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
LMArena Longer Query14791482

Writing & Preference GLM-5.3 leads

Claude Sonnet 4.6: 70.2 (#22), GLM-5.3: 75.7 (#6)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4.6GLM-5.3
LMArena Text14581471
LMArena Creative Writing14351457
EQ-Bench Creative Writing18102075
LMArena Multi-Turn14641472
EQ-Bench 41207—

Frequently asked questions

Is Claude Sonnet 4.6 better than GLM-5.3?

GLM-5.3 is the stronger model overall, scoring 54.8 to 50.3 on the Noometry Index.

Which is cheaper, Claude Sonnet 4.6 or GLM-5.3?

GLM-5.3 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; Claude Sonnet 4.6 lists at $3 and $15.

Is Claude Sonnet 4.6 or GLM-5.3 better for coding?

GLM-5.3 scores higher on coding benchmarks: 59.5 versus 46.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Claude Sonnet 4.6 and GLM-5.3 share?

37 benchmarks have published results for both models. Claude Sonnet 4.6 has 57 scored results on Noometry and GLM-5.3 has 42.

Related comparisons

Go deeper