OpenAI, proprietary

o4-mini

o4-mini by OpenAI ranks 132nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.6. Its strongest category is long context, where it ranks 33rd. API pricing starts at $1.10 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#132 of 354
Index score
41.6
Evidence
Confirmed 60 results
Provider
OpenAI
Released
April 16, 2025
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
100K
Input price
$1.10 / M
Output price
$4.40 / M
Blended price
$1.93 / M
Output speed
6 tokens/s Kagi
Value
#150 of 219
Knowledge cutoff
May 2024
Input
text, image

Category scores

Each category score combines every public result we have in that category.

o4-mini category scores
  1. Coding 40.9
  2. Agentic & Tool Use 32.6
  3. Reasoning 24.6
  4. Math 40.8
  5. Knowledge 43.6
  6. Multimodal 40.2
  7. Multilingual 47.0
  8. Instruction Following 75.2
  9. Long Context 45.5
  10. Writing & Preference 54.0
o4-mini category ranks
CategoryScoreRankResults
Coding40.9#1276
Agentic & Tool Use32.6#612
Reasoning24.6#16211
Math40.8#896
Knowledge43.6#918
Multimodal40.2#493
Multilingual47.0#1541
Instruction Following75.2#682
Long Context45.5#332
Writing & Preference54.0#1525

Strengths and weaknesses

Categories where o4-mini places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

o4-mini: strongest categories
CategoryScorevs medianRank
Long Context45.5+4.6#33 of 296, top 12%
Instruction Following75.2+3.9#68 of 305, top 23%
Math40.8+4.3#89 of 327, top 28%

Weakest categories

o4-mini: weakest categories
CategoryScorevs medianRank
Multilingual47.0−0.4#154 of 297, top 52%
Writing & Preference54.0+0.2#152 of 312, top 49%
Reasoning24.6+0.9#162 of 350, top 47%

Closest competitors

The models ranked just above and below o4-mini. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to o4-mini
ModelRankScoreBlended $/MSpeed
GPT-5 Mini#12841.8$0.693Compare
ERNIE 5.0 0110#12941.8——Compare
Granite 4.2 30b#13041.8——Compare
Muse Glimmer#13141.7——Compare
Gemini 3.5 Flash Lite#13341.5$0.85—Compare
Grok 4.1#13441.5——Compare
GLM-4.6#13541.4$112Compare
Grok 4.1 Fast#13641.4$0.28—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

o4-mini Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)45%#29 of 39, top 75%SWE-bench2025-07-26
Aider Polyglot72%#8 of 44, top 19%highEpoch AI
GSO3.6%#28 of 31, top 91%highEpoch AI
WeirdML52.6%#43 of 119, top 37%highEpoch AI
LMArena Coding1368#158 of 294, top 54%LMArena2026-10-08
CadEval62%#3 of 14, top 22%mediumEpoch AI
ALE-Bench826.17#54 of 105, top 52%highEpoch AI
AlgoTune1.72#6 of 18, top 34%highEpoch AI

Agentic & Tool Use

o4-mini Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard53.2%#16 of 49, top 33%fcBerkeley Function Calling Leaderboard
GDPval25.3%#8 of 11, top 73%highEpoch AI
METR Time Horizons63.9%#16 of 32, top 50%mediumEpoch AI

Reasoning

o4-mini Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-26.1%#52 of 83, top 63%highEpoch AI
ARC-AGI-21.7%lowEpoch AI
ARC-AGI-22.4%mediumEpoch AI
SimpleBench38.7%#55 of 77, top 72%highEpoch AI
Kagi LLM Benchmark67.6%#28 of 99, top 29%Kagi LLM Benchmark
ARC-AGI-158.7%#52 of 83, top 63%highEpoch AI
ARC-AGI-121.3%lowEpoch AI
ARC-AGI-141.8%mediumEpoch AI
CritPt0.6%#88 of 134, top 66%highEpoch AI
Chess Puzzles26%#41 of 129, top 32%highEpoch AI2025-12-08
Chess Puzzles14%lowEpoch AI2026-07-11
Chess Puzzles20%mediumEpoch AI2026-08-07
EnigmaEval9.2%#15 of 38, top 40%highEpoch AI
EnigmaEval6.8%mediumEpoch AI
LMArena Hard Prompts1351#159 of 297, top 54%LMArena2026-10-08
Mystery Game Puzzles5%#71 of 74, top 96%highEpoch AI2026-08-27
DTBench77.6%#80 of 151, top 53%highEpoch AI
LMCA26.5%#83 of 125, top 67%highEpoch AI
Epoch Capabilities Index145.64#80 of 213, top 38%Epoch AI2025-04-16
ForecastBench61.8#6 of 72, top 9%Epoch AI

Math

o4-mini Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)36.1%#55 of 81, top 68%highEpoch AI2026-06-11
FrontierMath (Tiers 1-3)16.1%lowEpoch AI2026-08-27
FrontierMath (Tiers 1-3)28.8%mediumEpoch AI2026-08-27
FrontierMath Tier 44.9%#56 of 63, top 89%highEpoch AI2026-06-11
OTIS Mock AIME 2024-202581.7%#78 of 173, top 46%highEpoch AI2025-04-16
OTIS Mock AIME 2024-202557.8%lowEpoch AI2026-07-13
OTIS Mock AIME 2024-202573.3%mediumEpoch AI2026-08-07
Omni-MATH72%#2 of 57, top 4%HELM Capabilities
LMArena Math1389#139 of 285, top 49%LMArena2026-10-08
MATH Level 597.8%#3 of 79, top 4%highEpoch AI2025-04-16
FrontierMath (Feb 2025 set)24.8%#25 of 68, top 37%highEpoch AI2025-11-13
FrontierMath (Feb 2025 set)10.7%lowEpoch AI2025-11-16
FrontierMath (Feb 2025 set)19%mediumEpoch AI2025-11-13
FrontierMath Tier 4 (v1)6.3%#25 of 55, top 46%highEpoch AI2025-07-01
FrontierMath Tier 4 (v1)2.1%mediumEpoch AI2025-08-07

Knowledge

o4-mini Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond79.6%#81 of 186, top 44%highEpoch AI2025-04-16
GPQA Diamond75.3%lowEpoch AI2026-07-11
GPQA Diamond77.8%mediumEpoch AI2026-08-07
Humanity's Last Exam18.1%#20 of 41, top 49%highEpoch AI
Humanity's Last Exam14.3%mediumEpoch AI
SimpleQA Verified18.8%highEpoch AI2026-08-27
SimpleQA Verified19.6%#66 of 77, top 86%lowEpoch AI2026-08-27
MMLU-Pro82%#12 of 58, top 21%HELM Capabilities
Confabulations (lower is better)15.8%#22 of 51, top 44%high reasoningLech Mazur benchmarks
Vectara Hallucination Rate (lower is better)18.6%#89 of 96, top 93%Vectara Hallucination Leaderboard
Vectara Hallucination Rate (lower is better)18.6%#89 of 96, top 93%Vectara Hallucination Leaderboard
GPQA (HELM)73.5%#6 of 57, top 11%HELM Capabilities
LMArena Expert1343#157 of 273, top 58%LMArena2026-10-08

Multimodal

o4-mini Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1194#82 of 122, top 68%LMArena2026-10-09
GeoBench64%#15 of 25, top 60%highEpoch AI
GeoBench64%mediumEpoch AI
VPCT57.5%#6 of 24, top 25%mediumEpoch AI

Multilingual

o4-mini Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1337#154 of 297, top 52%LMArena2026-10-08
LMArena Chinese1354#166 of 285, top 59%LMArena2026-10-08
LMArena French1364#139 of 223, top 63%LMArena2026-10-08
LMArena German1336#141 of 231, top 62%LMArena2026-10-08
LMArena Japanese1308#118 of 211, top 56%LMArena2026-10-08
LMArena Korean1312#124 of 213, top 59%LMArena2026-10-08
LMArena Russian1334#156 of 283, top 56%LMArena2026-10-08
LMArena Spanish1347#147 of 226, top 66%LMArena2026-10-08

Instruction Following

o4-mini Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval92.8%#5 of 57, top 9%HELM Capabilities
LMArena Instruction Following1321#162 of 298, top 55%LMArena2026-10-08

Long Context

o4-mini Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench77.8%#13 of 47, top 28%mediumEpoch AI
LMArena Longer Query1315#176 of 291, top 61%LMArena2026-10-08

Writing & Preference

o4-mini Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1353#156 of 297, top 53%LMArena2026-10-08
LMArena Creative Writing1294#170 of 295, top 58%LMArena2026-10-08
Short-Story Creative Writing75%#25 of 39, top 65%mediumEpoch AI
WildBench85.4%#11 of 57, top 20%HELM Capabilities
LMArena Multi-Turn1350#154 of 295, top 53%LMArena2026-10-08

API pricing by provider

o4-mini API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.10$4.40$0.282026-10-10
openai$1.10$4.40$0.282026-10-10
openrouter$1.10$4.40$0.282026-10-10

Compare o4-mini

Other OpenAI models

Frequently asked questions

How good is o4-mini?

o4-mini by OpenAI ranks 132nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.6. Its strongest category is long context, where it ranks 33rd. API pricing starts at $1.10 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.

How much does o4-mini cost?

o4-mini costs $1.10 per million input tokens and $4.40 per million output tokens on OpenAI's own API, with cached input at $0.28.

What is o4-mini's context window?

o4-mini accepts up to 200K tokens of input and can write up to 100K tokens in one response.

Is o4-mini open source?

No. o4-mini is proprietary and available only through OpenAI's API and partner platforms.

How fast is o4-mini?

o4-mini generated about 6 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are o4-mini's strengths and weaknesses?

Relative to other ranked models, o4-mini places best in long context, instruction following, math and lowest in multilingual, writing & preference, reasoning.

What is o4-mini best at?

Its best category is long context, where it ranks 33rd on Noometry.