Model comparison

Mercury 2 vs Trinity Large Thinking

Mercury 2 and Trinity Large Thinking score almost the same on the Noometry Index (39.1 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Mercury 2 Inception

39.1

Rank #175 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mercury 2 scores higher in 3 categories and Trinity Large Thinking in 4 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury 2 leads 23.8 to 16.9.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 12.3% for Mercury 2 and 6.9% for Trinity Large Thinking.
  • Both cost about the same: $0.25 input and $0.75 output per million tokens.
  • Trinity Large Thinking accepts more context: 262K tokens versus 128K.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Mercury 2 and Trinity Large Thinking specifications
Mercury 2Trinity Large Thinking
ProviderInceptionArcee AI
Noometry Index39.138.6
Released2026-02-202026-04-01
WeightsProprietaryOpen
Context window128K262K
Max output50K80K
Input $ / M tokens$0.25$0.25
Output $ / M tokens$0.75$0.80
Results tracked1724

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury 2: 33.5 (#255), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMercury 2Trinity Large Thinking
LMArena WebDev11711238
SciCode38.7%36.1%
LMArena Coding13911381
WeirdML43.2%—
ALE-Bench785.58—

Reasoning Mercury 2 leads

Mercury 2: 23.8 (#170), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMercury 2Trinity Large Thinking
CritPt0.8%0.9%
LMArena Hard Prompts13621350
NYT Connections (extended)—16.5%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Not comparable

Mercury 2: —, Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMercury 2Trinity Large Thinking
LMArena Math—1366

Knowledge Trinity Large Thinking leads

Mercury 2: 36.2 (#172), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMercury 2Trinity Large Thinking
Vectara Hallucination Rate12.3%6.9%
LMArena Expert13581360

Multilingual Too close to call

Mercury 2: 46.6 (#157), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMercury 2Trinity Large Thinking
LMArena Non-English13311325
LMArena Chinese14171373
LMArena Russian13041337
LMArena French—1374
LMArena German—1356
LMArena Japanese—1311
LMArena Korean—1306
LMArena Spanish—1357

Instruction Following Too close to call

Mercury 2: 70.2 (#165), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMercury 2Trinity Large Thinking
LMArena Instruction Following13291334

Long Context Too close to call

Mercury 2: 40.5 (#154), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMercury 2Trinity Large Thinking
LMArena Longer Query13301355

Writing & Preference Too close to call

Mercury 2: 53.8 (#155), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMercury 2Trinity Large Thinking
LMArena Text13551340
LMArena Creative Writing12891320
LMArena Multi-Turn13581342

Frequently asked questions

Is Mercury 2 better than Trinity Large Thinking?

Mercury 2 and Trinity Large Thinking score almost the same on the Noometry Index (39.1 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, Mercury 2 or Trinity Large Thinking?

Mercury 2 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; Trinity Large Thinking lists at $0.25 and $0.80.

Is Mercury 2 or Trinity Large Thinking better for coding?

They score almost the same on coding (33.5 vs 34.1); test both on your own repository before choosing.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 128K.

How many benchmarks do Mercury 2 and Trinity Large Thinking share?

15 benchmarks have published results for both models. Mercury 2 has 17 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper