Model comparison

Mistral Large 3 vs Trinity Large Thinking

Mistral Large 3 and Trinity Large Thinking score almost the same on the Noometry Index (39.1 vs 38.6), so choose on price, context window or the category you care about most.

Last verified . 21 shared benchmarks.

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Mistral Large 3 scores higher in 6 categories and Trinity Large Thinking in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral Large 3 leads 52.5 to 46.2.
  • The biggest single-benchmark swing is Thematic Generalization: 23% for Mistral Large 3 and 41.6% for Trinity Large Thinking.
  • Both cost about the same: $0.25 input and $0.75 output per million tokens.

Side by side

Mistral Large 3 and Trinity Large Thinking specifications
Mistral Large 3Trinity Large Thinking
ProviderMistral AIArcee AI
Noometry Index39.138.6
Released2025-12-022026-04-01
WeightsOpenOpen
Context window262K262K
Max output8K80K
Input $ / M tokens$0.25$0.25
Output $ / M tokens$0.75$0.80
Results tracked2424

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Large 3: 34.4 (#237), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
LMArena WebDev12301238
LMArena Coding14481381
SciCode—36.1%

Reasoning Trinity Large Thinking leads

Mistral Large 3: 15.2 (#319), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
NYT Connections (extended)7.5%16.5%
Thematic Generalization23%41.6%
LMArena Hard Prompts14291350
Kagi LLM Benchmark50.9%—
CritPt—0.9%
Surface Evolver Bench—15.6%

Math Mistral Large 3 leads

Mistral Large 3: 38.7 (#129), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
LMArena Math14141366

Knowledge Trinity Large Thinking leads

Mistral Large 3: 36.0 (#177), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
Vectara Hallucination Rate14.5%6.9%
LMArena Expert14211360

Multimodal Not comparable

Mistral Large 3: 38.2 (#66), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
LMArena Vision1221—

Multilingual Mistral Large 3 leads

Mistral Large 3: 52.5 (#84), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
LMArena Non-English14131325
LMArena Chinese14471373
LMArena French14551374
LMArena German14371356
LMArena Japanese13941311
LMArena Korean13841306
LMArena Russian14111337
LMArena Spanish14401357

Instruction Following Mistral Large 3 leads

Mistral Large 3: 74.0 (#108), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
LMArena Instruction Following14031334

Long Context Mistral Large 3 leads

Mistral Large 3: 43.1 (#105), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
LMArena Longer Query14131355

Writing & Preference Mistral Large 3 leads

Mistral Large 3: 60.0 (#101), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMistral Large 3Trinity Large Thinking
LMArena Text14281340
LMArena Creative Writing13861320
LMArena Multi-Turn14291342
EQ-Bench Creative Writing1412—

Frequently asked questions

Is Mistral Large 3 better than Trinity Large Thinking?

Mistral Large 3 and Trinity Large Thinking score almost the same on the Noometry Index (39.1 vs 38.6), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Large 3 or Trinity Large Thinking?

Mistral Large 3 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; Trinity Large Thinking lists at $0.25 and $0.80.

Is Mistral Large 3 or Trinity Large Thinking better for coding?

They score almost the same on coding (34.4 vs 34.1); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Mistral Large 3 and Trinity Large Thinking share?

21 benchmarks have published results for both models. Mistral Large 3 has 24 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper