Model comparison
Grok 4.3 vs MiniMax-M3
Grok 4.3 and MiniMax-M3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.
Last verified . 33 shared benchmarks.
Summary
- They share 33 benchmarks with published results for both. Grok 4.3 scores higher in 3 categories and MiniMax-M3 in 7 categories; 9 gaps are clear of the uncertainty.
- The widest gap is in multimodal, where MiniMax-M3 leads 40.2 to 31.6.
- The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 93.3% for Grok 4.3 and 71.1% for MiniMax-M3.
- MiniMax-M3 is cheaper at $0.30 / $1.20 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
- MiniMax-M3 has downloadable open weights; the other is API-only.
Side by side
| Grok 4.3 | MiniMax-M3 | |
|---|---|---|
| Provider | xAI | MiniMax |
| Noometry Index | 43.8 | 43.8 |
| Released | 2026-04-17 | 2026-06-01 |
| Weights | Proprietary | Open |
| Context window | 1M | 1M |
| Max output | 30K | 512K |
| Input $ / M tokens | $1.25 | $0.30 |
| Output $ / M tokens | $2.50 | $1.20 |
| Results tracked | 40 | 41 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Too close to call
Grok 4.3: 41.6 (#121), MiniMax-M3: 41.8 (#118)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| LMArena WebDev | 1357 | 1482 |
| SciCode | 47.3% | 47.1% |
| LMArena Coding | 1415 | 1469 |
| ALE-Bench | 944.17 | 640.02 |
| FrontierCode | — | 14.7% |
| WeirdML | 49.9% | — |
Agentic & Tool Use Grok 4.3 leads
Grok 4.3: 27.7 (#99), MiniMax-M3: 22.6 (#130)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| Vending-Bench 2 | 35.26 | 2,158 |
| APEX-Agents | — | 37.7% |
| OSWorld 2.0 | — | 4.6% |
| GBAEval | — | 0.9% |
| GDP.pdf | 8% | — |
| LMArena Search | 1165 | — |
Reasoning Grok 4.3 leads
Grok 4.3: 35.9 (#68), MiniMax-M3: 30.1 (#87)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| NYT Connections (extended) | 55.2% | 65.1% |
| CritPt | 8% | 3.7% |
| Chess Puzzles | 25% | 14% |
| LMArena Hard Prompts | 1396 | 1447 |
| DTBench | 90.7% | 78.9% |
| LMCA | 38.3% | 33.7% |
| Epoch Capabilities Index | 149.16 | 146.95 |
| ForecastBench | 60.3 | 61.4 |
| SimpleBench | — | 45.8% |
| Mystery Game Puzzles | — | 8% |
| Surface Evolver Bench | — | 55% |
Math Grok 4.3 leads
Grok 4.3: 46.0 (#74), MiniMax-M3: 40.0 (#95)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 93.3% | 71.1% |
| ProofBench | 11% | 18% |
| LMArena Math | 1388 | 1429 |
| FrontierMath (Tiers 1-3) | 42.8% | — |
| FrontierMath Tier 4 | 14.6% | — |
Knowledge MiniMax-M3 leads
Grok 4.3: 52.5 (#62), MiniMax-M3: 58.4 (#35)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| GPQA Diamond | 88.8% | 90.9% |
| LMArena Expert | 1385 | 1461 |
| SimpleQA Verified | 33.2% | — |
Multimodal MiniMax-M3 leads
Grok 4.3: 31.6 (#104), MiniMax-M3: 40.2 (#51)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| LMArena Vision | 1229 | 1253 |
| Blueprint-Bench 2 | 0% | — |
| LMArena Document | — | 1435 |
Multilingual MiniMax-M3 leads
Grok 4.3: 50.5 (#120), MiniMax-M3: 53.0 (#75)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| LMArena Non-English | 1385 | 1420 |
| LMArena Chinese | 1422 | 1463 |
| LMArena French | 1412 | 1447 |
| LMArena German | 1395 | 1426 |
| LMArena Japanese | 1379 | 1381 |
| LMArena Korean | 1356 | 1372 |
| LMArena Russian | 1399 | 1428 |
| LMArena Spanish | 1398 | 1432 |
Instruction Following MiniMax-M3 leads
Grok 4.3: 72.1 (#140), MiniMax-M3: 75.5 (#62)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| LMArena Instruction Following | 1366 | 1433 |
Long Context MiniMax-M3 leads
Grok 4.3: 42.5 (#123), MiniMax-M3: 44.2 (#72)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| LMArena Longer Query | 1393 | 1445 |
Writing & Preference MiniMax-M3 leads
Grok 4.3: 58.5 (#118), MiniMax-M3: 62.1 (#83)
| Benchmark | Grok 4.3 | MiniMax-M3 |
|---|---|---|
| LMArena Text | 1397 | 1433 |
| LMArena Creative Writing | 1380 | 1404 |
| EQ-Bench 4 | 1075 | 1150 |
| LMArena Multi-Turn | 1406 | 1442 |
Frequently asked questions
Is Grok 4.3 better than MiniMax-M3?
Grok 4.3 and MiniMax-M3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.
Which is cheaper, Grok 4.3 or MiniMax-M3?
MiniMax-M3 is cheaper. It lists at $0.30 per million input tokens and $1.20 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.
Is Grok 4.3 or MiniMax-M3 better for coding?
They score almost the same on coding (41.6 vs 41.8); test both on your own repository before choosing.
Which has the bigger context window?
Both accept 1M tokens.
How many benchmarks do Grok 4.3 and MiniMax-M3 share?
33 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and MiniMax-M3 has 41.