Model comparison

GLM-5 vs GPT-6 Astra

GPT-6 Astra is the stronger model overall, scoring 70.8 to 46.1 on the Noometry Index. GLM-5 costs 13× less per token, which makes it the better buy when GPT-6 Astra's lead doesn't matter for your workload.

Last verified . 30 shared benchmarks.

GLM-5 Z.ai (Zhipu)

46.1

Rank #66 Confirmed

GPT-6 Astra OpenAI

70.8

Rank #1 Confirmed

Summary

  • They share 30 benchmarks with published results for both. GLM-5 scores higher in 2 categories and GPT-6 Astra in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-6 Astra leads 85.1 to 27.6.
  • The biggest single-benchmark swing is ARC-AGI-2: 4.9% for GLM-5 and 95% for GPT-6 Astra.
  • GLM-5 is cheaper at $1 / $3.20 per million input/output tokens, against $10 / $50 for GPT-6 Astra.
  • GPT-6 Astra accepts more context: 1.05M tokens versus 205K.
  • GLM-5 has downloadable open weights; the other is API-only.

Side by side

GLM-5 and GPT-6 Astra specifications
GLM-5GPT-6 Astra
ProviderZ.ai (Zhipu)OpenAI
Noometry Index46.170.8
Released2026-02-112026-09-03
WeightsOpenProprietary
Context window205K1.05M
Max output131K128K
Input $ / M tokens$1$10
Output $ / M tokens$3.20$50
Results tracked4556

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Astra leads

GLM-5: 49.0 (#52), GPT-6 Astra: 73.7 (#2)

Coding benchmarks
BenchmarkGLM-5GPT-6 Astra
LMArena WebDev14341786
WeirdML48.2%93.6%
LMArena Coding14611487
ALE-Bench765.622,951
SWE-bench Verified72.1%—
DeepSWE—74.1%
FrontierCode—53.3%
SWE-bench Verified (bash only)72.8%—
SWE-bench Multilingual69.7%—
FrontierSWE—65.5%
SciCode—56.5%
GSO—79.4%
MirrorCode—46.7%

Agentic & Tool Use GPT-6 Astra leads

GLM-5: 31.1 (#71), GPT-6 Astra: 52.9 (#3)

Agentic & Tool Use benchmarks
BenchmarkGLM-5GPT-6 Astra
Vending-Bench 24,43215,515
Terminal-Bench52.4%—
APEX-Agents—64.7%
Remote Labor Index—20.8%
τ²-bench Airline82.5%—
τ²-bench Banking9.8%—
τ²-bench Retail73.7%—
τ²-bench Telecom86.8%—
BALROG—68.3%
GDP.pdf—34.2%

Reasoning GPT-6 Astra leads

GLM-5: 27.6 (#116), GPT-6 Astra: 85.1 (#1)

Reasoning benchmarks
BenchmarkGLM-5GPT-6 Astra
ARC-AGI-24.9%95%
NYT Connections (extended)74.8%98.1%
ARC-AGI-144.7%98.5%
Chess Puzzles10%72%
LMArena Hard Prompts14521462
Epoch Capabilities Index145.83166.45
SimpleBench53.2%—
Kagi LLM Benchmark75%—
CritPt—31.7%
EBR-Bench—76.2%
Mystery Game Puzzles—84%
DTBench—97.3%
LMCA—64.4%
Bench to the Future 3—0.14
ForecastBench61—

Math GPT-6 Astra leads

GLM-5: 46.4 (#71), GPT-6 Astra: 93.5 (#2)

Knowledge GPT-6 Astra leads

GLM-5: 52.3 (#64), GPT-6 Astra: 75.3 (#1)

Knowledge benchmarks
BenchmarkGLM-5GPT-6 Astra
GPQA Diamond87.8%95.8%
Vectara Hallucination Rate10.1%8.7%
LMArena Expert14541483
Humanity's Last Exam—54.8%
SimpleQA Verified—75.6%

Multimodal Not comparable

GLM-5: —, GPT-6 Astra: 55.0 (#3)

Multimodal benchmarks
BenchmarkGLM-5GPT-6 Astra
LMArena Vision—1281
Blueprint-Bench 2—49.7%
Furniture Assembly—80%
LMArena Document—1468

Multilingual Too close to call

GLM-5: 53.7 (#58), GPT-6 Astra: 53.7 (#61)

Multilingual benchmarks
BenchmarkGLM-5GPT-6 Astra
LMArena Non-English14301430
LMArena Chinese15111484
LMArena French14551456
LMArena German14451440
LMArena Japanese14161379
LMArena Korean14231426
LMArena Russian14361436
LMArena Spanish14541407

Instruction Following GPT-6 Astra leads

GLM-5: 75.2 (#67), GPT-6 Astra: 76.3 (#44)

Instruction Following benchmarks
BenchmarkGLM-5GPT-6 Astra
LMArena Instruction Following14281450

Long Context Too close to call

GLM-5: 44.7 (#60), GPT-6 Astra: 44.5 (#62)

Long Context benchmarks
BenchmarkGLM-5GPT-6 Astra
LMArena Longer Query14461456
CL-bench18.7%—

Writing & Preference GPT-6 Astra leads

GLM-5: 66.0 (#38), GPT-6 Astra: 75.3 (#7)

Writing & Preference benchmarks
BenchmarkGLM-5GPT-6 Astra
LMArena Text14461441
LMArena Creative Writing14391418
EQ-Bench Creative Writing16012173
LMArena Multi-Turn14561448

Frequently asked questions

Is GLM-5 better than GPT-6 Astra?

GPT-6 Astra is the stronger model overall, scoring 70.8 to 46.1 on the Noometry Index. GLM-5 costs 13× less per token, which makes it the better buy when GPT-6 Astra's lead doesn't matter for your workload.

Which is cheaper, GLM-5 or GPT-6 Astra?

GLM-5 is cheaper. It lists at $1 per million input tokens and $3.20 per million output tokens; GPT-6 Astra lists at $10 and $50.

Is GLM-5 or GPT-6 Astra better for coding?

GPT-6 Astra scores higher on coding benchmarks: 73.7 versus 49.0 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Astra does, with 1.05M tokens against 205K.

How many benchmarks do GLM-5 and GPT-6 Astra share?

30 benchmarks have published results for both models. GLM-5 has 45 scored results on Noometry and GPT-6 Astra has 56.

Related comparisons

Go deeper