Model comparison

Chatgpt 4o Latest 20250326 vs MiniMax-M3

Chatgpt 4o Latest 20250326 and MiniMax-M3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

MiniMax-M3 MiniMax

43.8

Rank #85 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 2 categories and MiniMax-M3 in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where MiniMax-M3 leads 58.4 to 39.2.
  • MiniMax-M3 has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and MiniMax-M3 specifications
Chatgpt 4o Latest 20250326MiniMax-M3
ProviderOpenAIMiniMax
Noometry Index43.843.8
Released—2026-06-01
WeightsProprietaryOpen
Context window—1M
Max output—512K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked2141

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Chatgpt 4o Latest 20250326: 41.6 (#122), MiniMax-M3: 41.8 (#118)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Coding14131469
FrontierCode—14.7%
LMArena WebDev—1482
SciCode—47.1%
ALE-Bench—640.02

Agentic & Tool Use Not comparable

Chatgpt 4o Latest 20250326: —, MiniMax-M3: 22.6 (#130)

Agentic & Tool Use benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
APEX-Agents—37.7%
OSWorld 2.0—4.6%
GBAEval—0.9%
Vending-Bench 2—2,158

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), MiniMax-M3: 30.1 (#87)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Hard Prompts14241447
SimpleBench—45.8%
Kagi LLM Benchmark75%—
NYT Connections (extended)—65.1%
CritPt—3.7%
Chess Puzzles—14%
Mystery Game Puzzles—8%
DTBench—78.9%
LMCA—33.7%
Surface Evolver Bench—55%
Epoch Capabilities Index—146.95
ForecastBench—61.4

Math MiniMax-M3 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), MiniMax-M3: 40.0 (#95)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Math14071429
OTIS Mock AIME 2024-2025—71.1%
ProofBench—18%

Knowledge MiniMax-M3 leads

Chatgpt 4o Latest 20250326: 39.2 (#137), MiniMax-M3: 58.4 (#35)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Expert14011461
GPQA Diamond—90.9%
Confabulations16.6%—

Multimodal Too close to call

Chatgpt 4o Latest 20250326: 39.6 (#58), MiniMax-M3: 40.2 (#51)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Vision12431253
LMArena Document—1435

Multilingual Too close to call

Chatgpt 4o Latest 20250326: 52.9 (#76), MiniMax-M3: 53.0 (#75)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Non-English14191420
LMArena Chinese14571463
LMArena French14461447
LMArena German14231426
LMArena Japanese14051381
LMArena Korean13961372
LMArena Russian14291428
LMArena Spanish14341432

Instruction Following MiniMax-M3 leads

Chatgpt 4o Latest 20250326: 74.0 (#107), MiniMax-M3: 75.5 (#62)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Instruction Following14031433

Long Context MiniMax-M3 leads

Chatgpt 4o Latest 20250326: 43.1 (#107), MiniMax-M3: 44.2 (#72)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Longer Query14131445

Writing & Preference Too close to call

Chatgpt 4o Latest 20250326: 62.6 (#72), MiniMax-M3: 62.1 (#83)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326MiniMax-M3
LMArena Text14291433
LMArena Creative Writing14051404
LMArena Multi-Turn14541442
EQ-Bench Creative Writing1501—
EQ-Bench 4—1150

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than MiniMax-M3?

Chatgpt 4o Latest 20250326 and MiniMax-M3 score almost the same on the Noometry Index (43.8 vs 43.8), so choose on price, context window or the category you care about most.

Is Chatgpt 4o Latest 20250326 or MiniMax-M3 better for coding?

They score almost the same on coding (41.6 vs 41.8); test both on your own repository before choosing.

How many benchmarks do Chatgpt 4o Latest 20250326 and MiniMax-M3 share?

18 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and MiniMax-M3 has 41.

Related comparisons

Go deeper