Model comparison

Grok 4.3 vs MiniMax-M3

Grok 4.3 and MiniMax-M3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 33 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

MiniMax-M3 MiniMax

43.8

Rank #85 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Grok 4.3 scores higher in 3 categories and MiniMax-M3 in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in multimodal, where MiniMax-M3 leads 40.2 to 31.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 93.3% for Grok 4.3 and 71.1% for MiniMax-M3.
  • MiniMax-M3 is cheaper at $0.30 / $1.20 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • MiniMax-M3 has downloadable open weights; the other is API-only.

Side by side

Grok 4.3 and MiniMax-M3 specifications
Grok 4.3MiniMax-M3
ProviderxAIMiniMax
Noometry Index43.843.8
Released2026-04-172026-06-01
WeightsProprietaryOpen
Context window1M1M
Max output30K512K
Input $ / M tokens$1.25$0.30
Output $ / M tokens$2.50$1.20
Results tracked4041

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.3: 41.6 (#121), MiniMax-M3: 41.8 (#118)

Coding benchmarks
BenchmarkGrok 4.3MiniMax-M3
LMArena WebDev13571482
SciCode47.3%47.1%
LMArena Coding14151469
ALE-Bench944.17640.02
FrontierCode—14.7%
WeirdML49.9%—

Agentic & Tool Use Grok 4.3 leads

Grok 4.3: 27.7 (#99), MiniMax-M3: 22.6 (#130)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3MiniMax-M3
Vending-Bench 235.262,158
APEX-Agents—37.7%
OSWorld 2.0—4.6%
GBAEval—0.9%
GDP.pdf8%—
LMArena Search1165—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), MiniMax-M3: 30.1 (#87)

Reasoning benchmarks
BenchmarkGrok 4.3MiniMax-M3
NYT Connections (extended)55.2%65.1%
CritPt8%3.7%
Chess Puzzles25%14%
LMArena Hard Prompts13961447
DTBench90.7%78.9%
LMCA38.3%33.7%
Epoch Capabilities Index149.16146.95
ForecastBench60.361.4
SimpleBench—45.8%
Mystery Game Puzzles—8%
Surface Evolver Bench—55%

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), MiniMax-M3: 40.0 (#95)

Math benchmarks
BenchmarkGrok 4.3MiniMax-M3
OTIS Mock AIME 2024-202593.3%71.1%
ProofBench11%18%
LMArena Math13881429
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—

Knowledge MiniMax-M3 leads

Grok 4.3: 52.5 (#62), MiniMax-M3: 58.4 (#35)

Knowledge benchmarks
BenchmarkGrok 4.3MiniMax-M3
GPQA Diamond88.8%90.9%
LMArena Expert13851461
SimpleQA Verified33.2%—

Multimodal MiniMax-M3 leads

Grok 4.3: 31.6 (#104), MiniMax-M3: 40.2 (#51)

Multimodal benchmarks
BenchmarkGrok 4.3MiniMax-M3
LMArena Vision12291253
Blueprint-Bench 20%—
LMArena Document—1435

Multilingual MiniMax-M3 leads

Grok 4.3: 50.5 (#120), MiniMax-M3: 53.0 (#75)

Multilingual benchmarks
BenchmarkGrok 4.3MiniMax-M3
LMArena Non-English13851420
LMArena Chinese14221463
LMArena French14121447
LMArena German13951426
LMArena Japanese13791381
LMArena Korean13561372
LMArena Russian13991428
LMArena Spanish13981432

Instruction Following MiniMax-M3 leads

Grok 4.3: 72.1 (#140), MiniMax-M3: 75.5 (#62)

Instruction Following benchmarks
BenchmarkGrok 4.3MiniMax-M3
LMArena Instruction Following13661433

Long Context MiniMax-M3 leads

Grok 4.3: 42.5 (#123), MiniMax-M3: 44.2 (#72)

Long Context benchmarks
BenchmarkGrok 4.3MiniMax-M3
LMArena Longer Query13931445

Writing & Preference MiniMax-M3 leads

Grok 4.3: 58.5 (#118), MiniMax-M3: 62.1 (#83)

Writing & Preference benchmarks
BenchmarkGrok 4.3MiniMax-M3
LMArena Text13971433
LMArena Creative Writing13801404
EQ-Bench 410751150
LMArena Multi-Turn14061442

Frequently asked questions

Is Grok 4.3 better than MiniMax-M3?

Grok 4.3 and MiniMax-M3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or MiniMax-M3?

MiniMax-M3 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is Grok 4.3 or MiniMax-M3 better for coding?

They score almost the same on coding (41.6 vs 41.8); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Grok 4.3 and MiniMax-M3 share?

33 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and MiniMax-M3 has 41.

Related comparisons

Go deeper