Model comparison

GLM-4.5-Air vs Trinity Large Thinking

GLM-4.5-Air and Trinity Large Thinking score almost the same on the Noometry Index (38.9 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 18 benchmarks with published results for both. GLM-4.5-Air scores higher in 4 categories and Trinity Large Thinking in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GLM-4.5-Air leads 24.1 to 16.9.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $0.20 / $1.10 for GLM-4.5-Air.
  • Trinity Large Thinking accepts more context: 262K tokens versus 131K.

Side by side

GLM-4.5-Air and Trinity Large Thinking specifications
GLM-4.5-AirTrinity Large Thinking
ProviderZ.ai (Zhipu)Arcee AI
Noometry Index38.938.6
Released2025-07-202026-04-01
WeightsOpenOpen
Context window131K262K
Max output98K80K
Input $ / M tokens$0.20$0.25
Output $ / M tokens$1.10$0.80
Results tracked2724

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-4.5-Air: 33.3 (#259), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
LMArena Coding13971381
LMArena WebDev—1238
SciCode—36.1%
GSO2.9%—

Reasoning GLM-4.5-Air leads

GLM-4.5-Air: 24.1 (#166), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
LMArena Hard Prompts13791350
Kagi LLM Benchmark43%—
NYT Connections (extended)—16.5%
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%
ForecastBench59.2—

Math Trinity Large Thinking leads

GLM-4.5-Air: 36.2 (#170), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
LMArena Math13961366
Omni-MATH39.1%—

Knowledge Trinity Large Thinking leads

GLM-4.5-Air: 35.0 (#191), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
Vectara Hallucination Rate9.3%6.9%
LMArena Expert13701360
Humanity's Last Exam8.1%—
MMLU-Pro76.2%—
GPQA (HELM)59.4%—

Multilingual GLM-4.5-Air leads

GLM-4.5-Air: 49.1 (#135), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
LMArena Non-English13661325
LMArena Chinese14261373
LMArena French13991374
LMArena German13771356
LMArena Japanese13481311
LMArena Korean13081306
LMArena Russian13731337
LMArena Spanish13861357

Instruction Following Too close to call

GLM-4.5-Air: 69.6 (#171), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
LMArena Instruction Following13541334
IFEval81.2%—

Long Context Too close to call

GLM-4.5-Air: 41.6 (#135), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
LMArena Longer Query13661355

Writing & Preference GLM-4.5-Air leads

GLM-4.5-Air: 55.9 (#139), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkGLM-4.5-AirTrinity Large Thinking
LMArena Text13841340
LMArena Creative Writing13431320
LMArena Multi-Turn13711342
WildBench78.9%—

Frequently asked questions

Is GLM-4.5-Air better than Trinity Large Thinking?

GLM-4.5-Air and Trinity Large Thinking score almost the same on the Noometry Index (38.9 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-4.5-Air or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; GLM-4.5-Air lists at $0.20 and $1.10.

Is GLM-4.5-Air or Trinity Large Thinking better for coding?

They score almost the same on coding (33.3 vs 34.1); test both on your own repository before choosing.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 131K.

How many benchmarks do GLM-4.5-Air and Trinity Large Thinking share?

18 benchmarks have published results for both models. GLM-4.5-Air has 27 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper