Model comparison

MiniMax-M2.7 vs MiniMax-M3

MiniMax-M3 is the stronger model overall, scoring 43.8 to 37.7 on the Noometry Index.

Last verified . 25 shared benchmarks.

MiniMax-M2.7 MiniMax

37.7

Rank #196 Confirmed

MiniMax-M3 MiniMax

43.8

Rank #85 Confirmed

Summary

  • They share 25 benchmarks with published results for both. MiniMax-M2.7 scores higher in 1 category and MiniMax-M3 in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where MiniMax-M3 leads 58.4 to 37.7.
  • The biggest single-benchmark swing is NYT Connections (extended): 24.7% for MiniMax-M2.7 and 65.1% for MiniMax-M3.
  • Both cost about the same: $0.30 input and $1.20 output per million tokens.
  • MiniMax-M3 accepts more context: 1M tokens versus 205K.

Side by side

MiniMax-M2.7 and MiniMax-M3 specifications
MiniMax-M2.7MiniMax-M3
ProviderMiniMaxMiniMax
Noometry Index37.743.8
Released2026-03-182026-06-01
WeightsOpenOpen
Context window205K1M
Max output131K512K
Input $ / M tokens$0.30$0.30
Output $ / M tokens$1.20$1.20
Results tracked3041

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

MiniMax-M2.7: 41.8 (#120), MiniMax-M3: 41.8 (#118)

Coding benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
LMArena WebDev13981482
SciCode47%47.1%
LMArena Coding14541469
ALE-Bench599.25640.02
FrontierCode—14.7%
WeirdML37%—

Agentic & Tool Use MiniMax-M2.7 leads

MiniMax-M2.7: 25.1 (#111), MiniMax-M3: 22.6 (#130)

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
GBAEval0%0.9%
Terminal-Bench45.1%—
APEX-Agents—37.7%
OSWorld 2.0—4.6%
ExploitBench13.3%—
Vending-Bench 2—2,158

Reasoning MiniMax-M3 leads

MiniMax-M2.7: 19.7 (#253), MiniMax-M3: 30.1 (#87)

Reasoning benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
NYT Connections (extended)24.7%65.1%
CritPt0.6%3.7%
LMArena Hard Prompts14221447
Epoch Capabilities Index145.85146.95
SimpleBench—45.8%
Chess Puzzles—14%
Thematic Generalization39.3%—
Mystery Game Puzzles—8%
DTBench—78.9%
LMCA—33.7%
Surface Evolver Bench—55%
ForecastBench—61.4

Math MiniMax-M3 leads

MiniMax-M2.7: 25.9 (#263), MiniMax-M3: 40.0 (#95)

Math benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
ProofBench3%18%
LMArena Math14201429
OTIS Mock AIME 2024-2025—71.1%

Knowledge MiniMax-M3 leads

MiniMax-M2.7: 37.7 (#152), MiniMax-M3: 58.4 (#35)

Knowledge benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
LMArena Expert14441461
GPQA Diamond—90.9%
Vectara Hallucination Rate12.9%—

Multimodal Not comparable

MiniMax-M2.7: —, MiniMax-M3: 40.2 (#51)

Multimodal benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
LMArena Vision—1253
LMArena Document—1435

Multilingual MiniMax-M3 leads

MiniMax-M2.7: 50.3 (#123), MiniMax-M3: 53.0 (#75)

Multilingual benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
LMArena Non-English13821420
LMArena Chinese14411463
LMArena French14211447
LMArena German13981426
LMArena Japanese12621381
LMArena Korean13131372
LMArena Russian13831428
LMArena Spanish14031432

Instruction Following MiniMax-M3 leads

MiniMax-M2.7: 74.1 (#103), MiniMax-M3: 75.5 (#62)

Instruction Following benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
LMArena Instruction Following14051433

Long Context Too close to call

MiniMax-M2.7: 43.3 (#99), MiniMax-M3: 44.2 (#72)

Long Context benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
LMArena Longer Query14191445

Writing & Preference MiniMax-M3 leads

MiniMax-M2.7: 58.9 (#112), MiniMax-M3: 62.1 (#83)

Writing & Preference benchmarks
BenchmarkMiniMax-M2.7MiniMax-M3
LMArena Text14051433
LMArena Creative Writing13541404
LMArena Multi-Turn14121442
EQ-Bench 4—1150

Frequently asked questions

Is MiniMax-M2.7 better than MiniMax-M3?

MiniMax-M3 is the stronger model overall, scoring 43.8 to 37.7 on the Noometry Index.

Which is cheaper, MiniMax-M2.7 or MiniMax-M3?

MiniMax-M3 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; MiniMax-M2.7 lists at $0.30 and $1.20.

Is MiniMax-M2.7 or MiniMax-M3 better for coding?

They score almost the same on coding (41.8 vs 41.8); test both on your own repository before choosing.

Which has the bigger context window?

MiniMax-M3 does, with 1M tokens against 205K.

How many benchmarks do MiniMax-M2.7 and MiniMax-M3 share?

25 benchmarks have published results for both models. MiniMax-M2.7 has 30 scored results on Noometry and MiniMax-M3 has 41.

Related comparisons

Go deeper