Model comparison

MiniMax-M2.5 vs Olmo 3.1 32b Think

MiniMax-M2.5 and Olmo 3.1 32b Think score almost the same on the Noometry Index (38.3 vs 37.9), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

MiniMax-M2.5 MiniMax

38.3

Rank #188 Confirmed

Summary

  • They share 15 benchmarks with published results for both. MiniMax-M2.5 scores higher in 5 categories and Olmo 3.1 32b Think in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where MiniMax-M2.5 leads 48.1 to 37.7.

Side by side

MiniMax-M2.5 and Olmo 3.1 32b Think specifications
MiniMax-M2.5Olmo 3.1 32b Think
ProviderMiniMaxAllen Institute for AI (Ai2)
Noometry Index38.337.9
Released2026-02-12—
WeightsOpenOpen
Context window205K—
Max output131K—
Input $ / M tokens$0.30—
Output $ / M tokens$1.20—
Results tracked3315

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.5 leads

MiniMax-M2.5: 48.1 (#58), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Coding13811291
SWE-bench Verified (bash only)75.8%—
LMArena WebDev1387—
SWE-bench Multilingual68.3%—
ALE-Bench618.17—

Agentic & Tool Use Not comparable

MiniMax-M2.5: 30.4 (#77), Olmo 3.1 32b Think: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
Terminal-Bench42.7%—
Vending-Bench 2-23.16—

Reasoning Olmo 3.1 32b Think leads

MiniMax-M2.5: 17.5 (#292), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Hard Prompts13721272
ARC-AGI-24.9%—
Kagi LLM Benchmark55.2%—
NYT Connections (extended)16.8%—
ARC-AGI-163.7%—
Epoch Capabilities Index146.68—

Math Olmo 3.1 32b Think leads

MiniMax-M2.5: 26.9 (#253), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Math13781305
ProofBench4%—

Knowledge MiniMax-M2.5 leads

MiniMax-M2.5: 39.2 (#135), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Expert13791295
Vectara Hallucination Rate9.1%—

Multilingual MiniMax-M2.5 leads

MiniMax-M2.5: 47.1 (#152), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Non-English13381209
LMArena Chinese13931242
LMArena French13621260
LMArena German13621262
LMArena Russian13581193
LMArena Spanish13541289
LMArena Japanese1171—
LMArena Korean1232—

Instruction Following MiniMax-M2.5 leads

MiniMax-M2.5: 71.5 (#148), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Instruction Following13531247

Long Context Olmo 3.1 32b Think leads

MiniMax-M2.5: 37.5 (#216), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Longer Query13661272
CL-bench11.4%—
CL-bench Life6.3%—

Writing & Preference MiniMax-M2.5 leads

MiniMax-M2.5: 53.9 (#153), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkMiniMax-M2.5Olmo 3.1 32b Think
LMArena Text13591272
LMArena Creative Writing13311226
LMArena Multi-Turn13641252
EQ-Bench Creative Writing1361—

Frequently asked questions

Is MiniMax-M2.5 better than Olmo 3.1 32b Think?

MiniMax-M2.5 and Olmo 3.1 32b Think score almost the same on the Noometry Index (38.3 vs 37.9), so choose on price, context window or the category you care about most.

Is MiniMax-M2.5 or Olmo 3.1 32b Think better for coding?

MiniMax-M2.5 scores higher on coding benchmarks: 48.1 versus 37.7 in the Noometry coding category.

How many benchmarks do MiniMax-M2.5 and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. MiniMax-M2.5 has 33 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper