Model comparison

GLM-4.7-Flash vs Trinity Large Thinking

GLM-4.7-Flash and Trinity Large Thinking score almost the same on the Noometry Index (38.8 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 17 benchmarks with published results for both. GLM-4.7-Flash scores higher in 3 categories and Trinity Large Thinking in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GLM-4.7-Flash leads 40.6 to 34.1.
  • GLM-4.7-Flash is cheaper at $0.06 / $0.40 per million input/output tokens, against $0.25 / $0.80 for Trinity Large Thinking.
  • Trinity Large Thinking accepts more context: 262K tokens versus 200K.

Side by side

GLM-4.7-Flash and Trinity Large Thinking specifications
GLM-4.7-FlashTrinity Large Thinking
ProviderZ.ai (Zhipu)Arcee AI
Noometry Index38.838.6
Released2026-01-192026-04-01
WeightsOpenOpen
Context window200K262K
Max output131K80K
Input $ / M tokens$0.06$0.25
Output $ / M tokens$0.40$0.80
Results tracked2124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

GLM-4.7-Flash: 40.6 (#135), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
LMArena Coding13831381
LMArena WebDev—1238
SciCode—36.1%

Reasoning GLM-4.7-Flash leads

GLM-4.7-Flash: 20.9 (#229), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
LMArena Hard Prompts13561350
NYT Connections (extended)—16.5%
CritPt—0.9%
Chess Puzzles0%—
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Trinity Large Thinking leads

GLM-4.7-Flash: 36.1 (#173), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
LMArena Math13551366
OTIS Mock AIME 2024-202558.3%—

Knowledge Trinity Large Thinking leads

GLM-4.7-Flash: 35.5 (#184), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
Vectara Hallucination Rate9.3%6.9%
LMArena Expert13571360
GPQA Diamond60.5%—

Multilingual Too close to call

GLM-4.7-Flash: 46.5 (#158), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
LMArena Non-English13301325
LMArena Chinese14031373
LMArena French13321374
LMArena German13371356
LMArena Korean12831306
LMArena Russian13321337
LMArena Spanish13501357
LMArena Japanese—1311

Instruction Following Too close to call

GLM-4.7-Flash: 70.1 (#167), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
LMArena Instruction Following13271334

Long Context Too close to call

GLM-4.7-Flash: 40.9 (#148), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
LMArena Longer Query13451355

Writing & Preference Trinity Large Thinking leads

GLM-4.7-Flash: 47.4 (#210), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashTrinity Large Thinking
LMArena Text13511340
LMArena Creative Writing12971320
LMArena Multi-Turn13421342
EQ-Bench Creative Writing1125—

Frequently asked questions

Is GLM-4.7-Flash better than Trinity Large Thinking?

GLM-4.7-Flash and Trinity Large Thinking score almost the same on the Noometry Index (38.8 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-4.7-Flash or Trinity Large Thinking?

GLM-4.7-Flash is cheaper. It lists at $0.06 per million input tokens and $0.40 per million output tokens; Trinity Large Thinking lists at $0.25 and $0.80.

Is GLM-4.7-Flash or Trinity Large Thinking better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 200K.

How many benchmarks do GLM-4.7-Flash and Trinity Large Thinking share?

17 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper