Model comparison

GLM-5 vs Hy4 preview

GLM-5 and Hy4 preview score almost the same on the Noometry Index (46.1 vs 45.3), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

GLM-5 Z.ai (Zhipu)

46.1

Rank #66 Confirmed

Hy4 preview Tencent

45.3

Rank #73 Reported

Summary

  • They share 2 benchmarks with published results for both. GLM-5 scores higher in 0 categories and Hy4 preview in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hy4 preview leads 55.7 to 46.4.
  • The biggest single-benchmark swing is NYT Connections (extended): 74.8% for GLM-5 and 68.2% for Hy4 preview.
  • Hy4 preview is cheaper at $0.75 / $2.25 per million input/output tokens, against $1 / $3.20 for GLM-5.
  • Hy4 preview accepts more context: 1.05M tokens versus 205K.

Side by side

GLM-5 and Hy4 preview specifications
GLM-5Hy4 preview
ProviderZ.ai (Zhipu)Tencent
Noometry Index46.145.3
Released2026-02-112026-08-28
WeightsOpenOpen
Context window205K1.05M
Max output131K64K
Input $ / M tokens$1$0.75
Output $ / M tokens$3.20$2.25
Results tracked453

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

GLM-5: 49.0 (#52), Hy4 preview: 51.6 (#38)

Coding benchmarks
BenchmarkGLM-5Hy4 preview
LMArena WebDev14341632
SWE-bench Verified72.1%—
SWE-bench Verified (bash only)72.8%—
SWE-bench Multilingual69.7%—
WeirdML48.2%—
LMArena Coding1461—
ALE-Bench765.62—

Agentic & Tool Use Not comparable

GLM-5: 31.1 (#71), Hy4 preview: —

Agentic & Tool Use benchmarks
BenchmarkGLM-5Hy4 preview
Terminal-Bench52.4%—
τ²-bench Airline82.5%—
τ²-bench Banking9.8%—
τ²-bench Retail73.7%—
τ²-bench Telecom86.8%—
Vending-Bench 24,432—

Reasoning Hy4 preview leads

GLM-5: 27.6 (#116), Hy4 preview: 31.9 (#79)

Reasoning benchmarks
BenchmarkGLM-5Hy4 preview
NYT Connections (extended)74.8%68.2%
ARC-AGI-24.9%—
SimpleBench53.2%—
Kagi LLM Benchmark75%—
ARC-AGI-144.7%—
Chess Puzzles10%—
LMArena Hard Prompts1452—
Epoch Capabilities Index145.83—
ForecastBench61—

Math Hy4 preview leads

GLM-5: 46.4 (#71), Hy4 preview: 55.7 (#42)

Knowledge Not comparable

GLM-5: 52.3 (#64), Hy4 preview: —

Knowledge benchmarks
BenchmarkGLM-5Hy4 preview
GPQA Diamond87.8%—
Vectara Hallucination Rate10.1%—
LMArena Expert1454—

Multilingual Not comparable

GLM-5: 53.7 (#58), Hy4 preview: —

Multilingual benchmarks
BenchmarkGLM-5Hy4 preview
LMArena Non-English1430—
LMArena Chinese1511—
LMArena French1455—
LMArena German1445—
LMArena Japanese1416—
LMArena Korean1423—
LMArena Russian1436—
LMArena Spanish1454—

Instruction Following Not comparable

GLM-5: 75.2 (#67), Hy4 preview: —

Instruction Following benchmarks
BenchmarkGLM-5Hy4 preview
LMArena Instruction Following1428—

Long Context Not comparable

GLM-5: 44.7 (#60), Hy4 preview: —

Long Context benchmarks
BenchmarkGLM-5Hy4 preview
CL-bench18.7%—
LMArena Longer Query1446—

Writing & Preference Not comparable

GLM-5: 66.0 (#38), Hy4 preview: —

Writing & Preference benchmarks
BenchmarkGLM-5Hy4 preview
LMArena Text1446—
LMArena Creative Writing1439—
EQ-Bench Creative Writing1601—
LMArena Multi-Turn1456—

Frequently asked questions

Is GLM-5 better than Hy4 preview?

GLM-5 and Hy4 preview score almost the same on the Noometry Index (46.1 vs 45.3), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5 or Hy4 preview?

Hy4 preview is cheaper. It lists at $0.75 per million input tokens and $2.25 per million output tokens; GLM-5 lists at $1 and $3.20.

Is GLM-5 or Hy4 preview better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 49.0 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 205K.

How many benchmarks do GLM-5 and Hy4 preview share?

2 benchmarks have published results for both models. GLM-5 has 45 scored results on Noometry and Hy4 preview has 3.

Related comparisons

Go deeper