Model comparison

MiniMax-M3 vs o4-mini

MiniMax-M3 is the stronger model overall, scoring 43.8 to 41.6 on the Noometry Index.

Last verified . 29 shared benchmarks.

MiniMax-M3 MiniMax

43.8

Rank #85 Confirmed

o4-mini OpenAI

41.6

Rank #132 Confirmed

Summary

  • They share 29 benchmarks with published results for both. MiniMax-M3 scores higher in 6 categories and o4-mini in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where MiniMax-M3 leads 58.4 to 43.6.
  • The biggest single-benchmark swing is Chess Puzzles: 14% for MiniMax-M3 and 26% for o4-mini.
  • MiniMax-M3 is cheaper at $0.30 / $1.20 per million input/output tokens, against $1.10 / $4.40 for o4-mini.
  • MiniMax-M3 accepts more context: 1M tokens versus 200K.
  • MiniMax-M3 has downloadable open weights; the other is API-only.

Side by side

MiniMax-M3 and o4-mini specifications
MiniMax-M3o4-mini
ProviderMiniMaxOpenAI
Noometry Index43.841.6
Released2026-06-012025-04-16
WeightsOpenProprietary
Context window1M200K
Max output512K100K
Input $ / M tokens$0.30$1.10
Output $ / M tokens$1.20$4.40
Results tracked4160

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

MiniMax-M3: 41.8 (#118), o4-mini: 40.9 (#127)

Coding benchmarks
BenchmarkMiniMax-M3o4-mini
LMArena Coding14691368
ALE-Bench640.02826.17
FrontierCode14.7%—
SWE-bench Verified (bash only)—45%
Aider Polyglot—72%
LMArena WebDev1482—
SciCode47.1%—
GSO—3.6%
WeirdML—52.6%
CadEval—62%
AlgoTune—1.72

Agentic & Tool Use o4-mini leads

MiniMax-M3: 22.6 (#130), o4-mini: 32.6 (#61)

Agentic & Tool Use benchmarks
BenchmarkMiniMax-M3o4-mini
APEX-Agents37.7%—
Berkeley Function Calling Leaderboard—53.2%
OSWorld 2.04.6%—
GDPval—25.3%
GBAEval0.9%—
METR Time Horizons—63.9%
Vending-Bench 22,158—

Reasoning MiniMax-M3 leads

MiniMax-M3: 30.1 (#87), o4-mini: 24.6 (#162)

Reasoning benchmarks
BenchmarkMiniMax-M3o4-mini
SimpleBench45.8%38.7%
CritPt3.7%0.6%
Chess Puzzles14%26%
LMArena Hard Prompts14471351
Mystery Game Puzzles8%5%
DTBench78.9%77.6%
LMCA33.7%26.5%
Epoch Capabilities Index146.95145.64
ForecastBench61.461.8
ARC-AGI-2—6.1%
Kagi LLM Benchmark—67.6%
NYT Connections (extended)65.1%—
ARC-AGI-1—58.7%
EnigmaEval—9.2%
Surface Evolver Bench55%—

Math Too close to call

MiniMax-M3: 40.0 (#95), o4-mini: 40.8 (#89)

Math benchmarks
BenchmarkMiniMax-M3o4-mini
OTIS Mock AIME 2024-202571.1%81.7%
LMArena Math14291389
FrontierMath (Tiers 1-3)—36.1%
FrontierMath Tier 4—4.9%
ProofBench18%—
Omni-MATH—72%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—24.8%
FrontierMath Tier 4 (v1)—6.3%

Knowledge MiniMax-M3 leads

MiniMax-M3: 58.4 (#35), o4-mini: 43.6 (#91)

Knowledge benchmarks
BenchmarkMiniMax-M3o4-mini
GPQA Diamond90.9%79.6%
LMArena Expert14611343
Humanity's Last Exam—18.1%
SimpleQA Verified—19.6%
MMLU-Pro—82%
Confabulations—15.8%
Vectara Hallucination Rate—18.6%
GPQA (HELM)—73.5%

Multimodal Too close to call

MiniMax-M3: 40.2 (#51), o4-mini: 40.2 (#49)

Multimodal benchmarks
BenchmarkMiniMax-M3o4-mini
LMArena Vision12531194
GeoBench—64%
VPCT—57.5%
LMArena Document1435—

Multilingual MiniMax-M3 leads

MiniMax-M3: 53.0 (#75), o4-mini: 47.0 (#154)

Multilingual benchmarks
BenchmarkMiniMax-M3o4-mini
LMArena Non-English14201337
LMArena Chinese14631354
LMArena French14471364
LMArena German14261336
LMArena Japanese13811308
LMArena Korean13721312
LMArena Russian14281334
LMArena Spanish14321347

Instruction Following Too close to call

MiniMax-M3: 75.5 (#62), o4-mini: 75.2 (#68)

Instruction Following benchmarks
BenchmarkMiniMax-M3o4-mini
LMArena Instruction Following14331321
IFEval—92.8%

Long Context o4-mini leads

MiniMax-M3: 44.2 (#72), o4-mini: 45.5 (#33)

Long Context benchmarks
BenchmarkMiniMax-M3o4-mini
LMArena Longer Query14451315
Fiction.LiveBench—77.8%

Writing & Preference MiniMax-M3 leads

MiniMax-M3: 62.1 (#83), o4-mini: 54.0 (#152)

Writing & Preference benchmarks
BenchmarkMiniMax-M3o4-mini
LMArena Text14331353
LMArena Creative Writing14041294
LMArena Multi-Turn14421350
Short-Story Creative Writing—75%
WildBench—85.4%
EQ-Bench 41150—

Frequently asked questions

Is MiniMax-M3 better than o4-mini?

MiniMax-M3 is the stronger model overall, scoring 43.8 to 41.6 on the Noometry Index.

Which is cheaper, MiniMax-M3 or o4-mini?

MiniMax-M3 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; o4-mini lists at $1.10 and $4.40.

Is MiniMax-M3 or o4-mini better for coding?

They score almost the same on coding (41.8 vs 40.9); test both on your own repository before choosing.

Which has the bigger context window?

MiniMax-M3 does, with 1M tokens against 200K.

How many benchmarks do MiniMax-M3 and o4-mini share?

29 benchmarks have published results for both models. MiniMax-M3 has 41 scored results on Noometry and o4-mini has 60.

Related comparisons

Go deeper