Model comparison

MiniMax-M2.7 vs Olmo 3.1 32b Think

MiniMax-M2.7 and Olmo 3.1 32b Think score almost the same on the Noometry Index (37.7 vs 37.9), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

Summary

  • They share 15 benchmarks with published results for both. MiniMax-M2.7 scores higher in 6 categories and Olmo 3.1 32b Think in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where MiniMax-M2.7 leads 58.9 to 46.2.

Side by side

MiniMax-M2.7 and Olmo 3.1 32b Think specifications
MiniMax-M2.7Olmo 3.1 32b Think
ProviderMiniMaxAllen Institute for AI (Ai2)
Noometry Index37.737.9
Released2026-03-18—
WeightsOpenOpen
Context window205K—
Max output131K—
Input $ / M tokens$0.30—
Output $ / M tokens$1.20—
Results tracked3015

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiniMax-M2.7 leads

MiniMax-M2.7: 41.8 (#120), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Coding14541291
LMArena WebDev1398—
SciCode47%—
WeirdML37%—
ALE-Bench599.25—

Agentic & Tool Use Not comparable

MiniMax-M2.7: 25.1 (#111), Olmo 3.1 32b Think: —

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
Terminal-Bench45.1%—
ExploitBench13.3%—
GBAEval0%—

Reasoning Olmo 3.1 32b Think leads

MiniMax-M2.7: 19.7 (#253), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Hard Prompts14221272
NYT Connections (extended)24.7%—
CritPt0.6%—
Thematic Generalization39.3%—
Epoch Capabilities Index145.85—

Math Olmo 3.1 32b Think leads

MiniMax-M2.7: 25.9 (#263), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Math14201305
ProofBench3%—

Knowledge MiniMax-M2.7 leads

MiniMax-M2.7: 37.7 (#152), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Expert14441295
Vectara Hallucination Rate12.9%—

Multilingual MiniMax-M2.7 leads

MiniMax-M2.7: 50.3 (#123), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Non-English13821209
LMArena Chinese14411242
LMArena French14211260
LMArena German13981262
LMArena Russian13831193
LMArena Spanish14031289
LMArena Japanese1262—
LMArena Korean1313—

Instruction Following MiniMax-M2.7 leads

MiniMax-M2.7: 74.1 (#103), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Instruction Following14051247

Long Context MiniMax-M2.7 leads

MiniMax-M2.7: 43.3 (#99), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Longer Query14191272

Writing & Preference MiniMax-M2.7 leads

MiniMax-M2.7: 58.9 (#112), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkMiniMax-M2.7Olmo 3.1 32b Think
LMArena Text14051272
LMArena Creative Writing13541226
LMArena Multi-Turn14121252

Frequently asked questions

Is MiniMax-M2.7 better than Olmo 3.1 32b Think?

MiniMax-M2.7 and Olmo 3.1 32b Think score almost the same on the Noometry Index (37.7 vs 37.9), so choose on price, context window or the category you care about most.

Is MiniMax-M2.7 or Olmo 3.1 32b Think better for coding?

MiniMax-M2.7 scores higher on coding benchmarks: 41.8 versus 37.7 in the Noometry coding category.

How many benchmarks do MiniMax-M2.7 and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. MiniMax-M2.7 has 30 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper