Model comparison

MiniMax-M2.5 vs Trinity Large Thinking

MiniMax-M2.5 and Trinity Large Thinking score almost the same on the Noometry Index (38.3 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

MiniMax-M2.5 MiniMax

38.3

Rank #188 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 20 benchmarks with published results for both. MiniMax-M2.5 scores higher in 5 categories and Trinity Large Thinking in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where MiniMax-M2.5 leads 48.1 to 34.1.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M2.5.
  • Trinity Large Thinking accepts more context: 262K tokens versus 205K.

Side by side

MiniMax-M2.5 and Trinity Large Thinking specifications
MiniMax-M2.5Trinity Large Thinking
ProviderMiniMaxArcee AI
Noometry Index38.338.6
Released2026-02-122026-04-01
WeightsOpenOpen
Context window205K262K
Max output131K80K
Input $ / M tokens$0.30$0.25
Output $ / M tokens$1.20$0.80
Results tracked3324

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.5 leads

MiniMax-M2.5: 48.1 (#58), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
LMArena WebDev13871238
LMArena Coding13811381
SWE-bench Verified (bash only)75.8%—
SWE-bench Multilingual68.3%—
SciCode—36.1%
ALE-Bench618.17—

Agentic & Tool Use Not comparable

MiniMax-M2.5: 30.4 (#77), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
Terminal-Bench42.7%—
Vending-Bench 2-23.16—

Reasoning Too close to call

MiniMax-M2.5: 17.5 (#292), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
NYT Connections (extended)16.8%16.5%
LMArena Hard Prompts13721350
ARC-AGI-24.9%—
Kagi LLM Benchmark55.2%—
ARC-AGI-163.7%—
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%
Epoch Capabilities Index146.68—

Math Trinity Large Thinking leads

MiniMax-M2.5: 26.9 (#253), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
LMArena Math13781366
ProofBench4%—

Knowledge Trinity Large Thinking leads

MiniMax-M2.5: 39.2 (#135), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
Vectara Hallucination Rate9.1%6.9%
LMArena Expert13791360

Multilingual Too close to call

MiniMax-M2.5: 47.1 (#152), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
LMArena Non-English13381325
LMArena Chinese13931373
LMArena French13621374
LMArena German13621356
LMArena Japanese11711311
LMArena Korean12321306
LMArena Russian13581337
LMArena Spanish13541357

Instruction Following MiniMax-M2.5 leads

MiniMax-M2.5: 71.5 (#148), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
LMArena Instruction Following13531334

Long Context Trinity Large Thinking leads

MiniMax-M2.5: 37.5 (#216), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
LMArena Longer Query13661355
CL-bench11.4%—
CL-bench Life6.3%—

Writing & Preference Too close to call

MiniMax-M2.5: 53.9 (#153), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMiniMax-M2.5Trinity Large Thinking
LMArena Text13591340
LMArena Creative Writing13311320
LMArena Multi-Turn13641342
EQ-Bench Creative Writing1361—

Frequently asked questions

Is MiniMax-M2.5 better than Trinity Large Thinking?

MiniMax-M2.5 and Trinity Large Thinking score almost the same on the Noometry Index (38.3 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, MiniMax-M2.5 or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; MiniMax-M2.5 lists at $0.30 and $1.20.

Is MiniMax-M2.5 or Trinity Large Thinking better for coding?

MiniMax-M2.5 scores higher on coding benchmarks: 48.1 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 205K.

How many benchmarks do MiniMax-M2.5 and Trinity Large Thinking share?

20 benchmarks have published results for both models. MiniMax-M2.5 has 33 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper