Model comparison

MiniMax-M2.1 vs Trinity Large Thinking

MiniMax-M2.1 and Trinity Large Thinking score almost the same on the Noometry Index (38.9 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 20 shared benchmarks.

MiniMax-M2.1 MiniMax

38.9

Rank #178 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 20 benchmarks with published results for both. MiniMax-M2.1 scores higher in 6 categories and Trinity Large Thinking in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where MiniMax-M2.1 leads 40.4 to 34.1.
  • The biggest single-benchmark swing is NYT Connections (extended): 11.2% for MiniMax-M2.1 and 16.5% for Trinity Large Thinking.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $0.30 / $1.20 for MiniMax-M2.1.
  • Trinity Large Thinking accepts more context: 262K tokens versus 205K.

Side by side

MiniMax-M2.1 and Trinity Large Thinking specifications
MiniMax-M2.1Trinity Large Thinking
ProviderMiniMaxArcee AI
Noometry Index38.938.6
Released2025-12-232026-04-01
WeightsOpenOpen
Context window205K262K
Max output131K80K
Input $ / M tokens$0.30$0.25
Output $ / M tokens$1.20$0.80
Results tracked2224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.1 leads

MiniMax-M2.1: 40.4 (#143), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
LMArena WebDev13841238
LMArena Coding14211381
SciCode—36.1%
ALE-Bench623.83—

Agentic & Tool Use Not comparable

MiniMax-M2.1: 27.9 (#98), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
Terminal-Bench36.6%—

Reasoning Too close to call

MiniMax-M2.1: 16.6 (#302), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
NYT Connections (extended)11.2%16.5%
LMArena Hard Prompts14111350
CritPt—0.9%
Thematic Generalization—41.6%
Surface Evolver Bench—15.6%

Math Too close to call

MiniMax-M2.1: 38.3 (#138), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
LMArena Math13971366

Knowledge Trinity Large Thinking leads

MiniMax-M2.1: 38.3 (#147), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
Vectara Hallucination Rate11.8%6.9%
LMArena Expert14311360

Multilingual MiniMax-M2.1 leads

MiniMax-M2.1: 50.0 (#128), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
LMArena Non-English13781325
LMArena Chinese14301373
LMArena French14041374
LMArena German13811356
LMArena Japanese12871311
LMArena Korean12981306
LMArena Russian13871337
LMArena Spanish13971357

Instruction Following MiniMax-M2.1 leads

MiniMax-M2.1: 73.8 (#112), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
LMArena Instruction Following14001334

Long Context MiniMax-M2.1 leads

MiniMax-M2.1: 43.2 (#101), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
LMArena Longer Query14161355

Writing & Preference MiniMax-M2.1 leads

MiniMax-M2.1: 58.3 (#120), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMiniMax-M2.1Trinity Large Thinking
LMArena Text13921340
LMArena Creative Writing13611320
LMArena Multi-Turn13961342

Frequently asked questions

Is MiniMax-M2.1 better than Trinity Large Thinking?

MiniMax-M2.1 and Trinity Large Thinking score almost the same on the Noometry Index (38.9 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, MiniMax-M2.1 or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; MiniMax-M2.1 lists at $0.30 and $1.20.

Is MiniMax-M2.1 or Trinity Large Thinking better for coding?

MiniMax-M2.1 scores higher on coding benchmarks: 40.4 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Trinity Large Thinking does, with 262K tokens against 205K.

How many benchmarks do MiniMax-M2.1 and Trinity Large Thinking share?

20 benchmarks have published results for both models. MiniMax-M2.1 has 22 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper