OpenAI, proprietary

o3-mini

o3-mini by OpenAI ranks 212th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.7. Its strongest category is instruction following, where it ranks 72nd. API pricing starts at $1.10 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.

Last verified

Specifications

Noometry rank
#212 of 354
Index score
36.7
Evidence
Confirmed 51 results
Provider
OpenAI
Released
December 20, 2024
Weights
Proprietary
Reasoning
Yes
Context window
200K
Max output
100K
Input price
$1.10 / M
Output price
$4.40 / M
Blended price
$1.93 / M
Output speed
Not measured
Value
#152 of 219
Knowledge cutoff
May 2024
Input
text

Category scores

Each category score combines every public result we have in that category.

o3-mini category scores
  1. Coding 40.8
  2. Agentic & Tool Use 29.6
  3. Reasoning 16.3
  4. Math 28.1
  5. Knowledge 38.3
  6. Multilingual 45.7
  7. Instruction Following 75.1
  8. Long Context 33.8
  9. Writing & Preference 50.3
o3-mini category ranks
CategoryScoreRankResults
Coding40.8#1327
Agentic & Tool Use29.6#841
Reasoning16.3#30511
Math28.1#2446
Knowledge38.3#1464
Multilingual45.7#1641
Instruction Following75.1#722
Long Context33.8#2562
Writing & Preference50.3#1825

Strengths and weaknesses

Categories where o3-mini places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

o3-mini: strongest categories
CategoryScorevs medianRank
Instruction Following75.1+3.8#72 of 305, top 24%
Coding40.8+2.1#132 of 340, top 39%
Knowledge38.3+1.0#146 of 314, top 47%

Weakest categories

o3-mini: weakest categories
CategoryScorevs medianRank
Reasoning16.3−7.3#305 of 350, top 88%
Long Context33.8−7.1#256 of 296, top 87%
Math28.1−8.5#244 of 327, top 75%

Closest competitors

The models ranked just above and below o3-mini. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to o3-mini
ModelRankScoreBlended $/MSpeed
GPT-4.5#20837.2——Compare
Yi-Lightning#20937.1——Compare
Qwen Plus#21037.1$0.6037Compare
Gemini 2.5 Flash-Lite#21137.0$0.18172Compare
Llama 3.1 Nemotron Ultra 253b v1#21336.7——Compare
Granite 4.0 H Small#21436.5——Compare
Command A#21536.5$4.3828Compare
Grok Build 0.1#21636.4$1.25—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

o3-mini Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot60.4%#13 of 44, top 30%highEpoch AI
Aider Polyglot53.8%mediumEpoch AI
SciCode39.8%#82 of 121, top 68%highEpoch AI
GSO1.3%#30 of 31, top 97%highEpoch AI
GSO1.3%lowEpoch AI
WeirdML43.7%#64 of 119, top 54%highEpoch AI
LiveBench Coding82.7%#2 of 39, top 6%highEpoch AI
LiveBench Coding61.5%lowEpoch AI
LiveBench Coding65.4%mediumEpoch AI
LMArena Coding1378#153 of 294, top 53%highLMArena2026-10-08
CadEval54%#6 of 14, top 43%mediumEpoch AI

Agentic & Tool Use

o3-mini Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Cybench22.5%#10 of 21, top 48%mediumEpoch AI

Reasoning

o3-mini Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-23%#63 of 83, top 76%highEpoch AI
ARC-AGI-20%lowEpoch AI
ARC-AGI-22.1%mediumEpoch AI
SimpleBench22.8%#69 of 77, top 90%highEpoch AI
ARC-AGI-134.5%#63 of 83, top 76%highEpoch AI
ARC-AGI-114.5%lowEpoch AI
ARC-AGI-122.3%mediumEpoch AI
CritPt0.3%#96 of 134, top 72%highEpoch AI
Chess Puzzles17%#66 of 129, top 52%highEpoch AI2025-12-08
Chess Puzzles6%lowEpoch AI2026-07-15
Chess Puzzles9%mediumEpoch AI2026-08-07
LiveBench Reasoning89.6%#4 of 39, top 11%highEpoch AI
LiveBench Reasoning69.8%lowEpoch AI
LiveBench Reasoning86.3%mediumEpoch AI
LMArena Hard Prompts1366#149 of 297, top 51%highLMArena2026-10-08
Mystery Game Puzzles7%#68 of 74, top 92%highEpoch AI2026-08-27
DTBench68.8%#95 of 151, top 63%highEpoch AI
LiveBench Data Analysis70.6%#4 of 39, top 11%highEpoch AI
LiveBench Data Analysis62%lowEpoch AI
LiveBench Data Analysis66.6%mediumEpoch AI
LMCA19%#94 of 125, top 76%highEpoch AI
Epoch Capabilities Index140.34#107 of 213, top 51%Epoch AI2025-01-31
ForecastBench59.6#37 of 72, top 52%Epoch AI
LiveBench75.9%#4 of 39, top 11%highEpoch AI
LiveBench62.5%lowEpoch AI
LiveBench70%mediumEpoch AI

Math

o3-mini Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)18.6%#72 of 81, top 89%highEpoch AI2026-06-11
FrontierMath (Tiers 1-3)3.9%lowEpoch AI2026-08-27
FrontierMath (Tiers 1-3)10.5%mediumEpoch AI2026-08-27
FrontierMath Tier 40%#63 of 63, top 100%highEpoch AI2026-06-11
OTIS Mock AIME 2024-202576.9%#84 of 173, top 49%highEpoch AI2025-02-27
OTIS Mock AIME 2024-202544.4%lowEpoch AI2026-07-20
OTIS Mock AIME 2024-202563.9%mediumEpoch AI2025-02-25
LiveBench Math77.3%#7 of 39, top 18%highEpoch AI
LiveBench Math63.1%lowEpoch AI
LiveBench Math72.4%mediumEpoch AI
LMArena Math1396#132 of 285, top 47%highLMArena2026-10-08
MATH Level 596.5%#8 of 79, top 11%highEpoch AI2025-02-13
MATH Level 595.2%mediumEpoch AI2025-01-31
FrontierMath (Feb 2025 set)12.4%#36 of 68, top 53%highEpoch AI2025-11-16
FrontierMath (Feb 2025 set)8.1%mediumEpoch AI2025-03-06
FrontierMath Tier 4 (v1)4.2%#34 of 55, top 62%highEpoch AI2025-07-01

Knowledge

o3-mini Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond77%#86 of 186, top 47%highEpoch AI2025-02-13
GPQA Diamond68.2%lowEpoch AI2026-07-20
GPQA Diamond74.3%mediumEpoch AI2025-01-31
SimpleQA Verified15.3%#69 of 77, top 90%highEpoch AI2026-08-31
Confabulations (lower is better)17.9%#28 of 51, top 55%medium reasoningLech Mazur benchmarks
LMArena Expert1364#145 of 273, top 54%highLMArena2026-10-08

Multilingual

o3-mini Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1319#164 of 297, top 56%highLMArena2026-10-08
LMArena Chinese1379#152 of 285, top 54%highLMArena2026-10-08
LMArena French1334#150 of 223, top 68%highLMArena2026-10-08
LMArena German1303#149 of 231, top 65%highLMArena2026-10-08
LMArena Japanese1286#129 of 211, top 62%highLMArena2026-10-08
LMArena Korean1314#121 of 213, top 57%highLMArena2026-10-08
LMArena Russian1304#175 of 283, top 62%highLMArena2026-10-08
LMArena Spanish1321#154 of 226, top 69%highLMArena2026-10-08

Instruction Following

o3-mini Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following84.4%#3 of 39, top 8%highEpoch AI
LiveBench Instruction Following80.1%lowEpoch AI
LiveBench Instruction Following83.2%mediumEpoch AI
LMArena Instruction Following1337#151 of 298, top 51%highLMArena2026-10-08

Long Context

o3-mini Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench50%#35 of 47, top 75%mediumEpoch AI
LMArena Longer Query1343#157 of 291, top 54%highLMArena2026-10-08

Writing & Preference

o3-mini Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1337#168 of 297, top 57%highLMArena2026-10-08
LMArena Creative Writing1286#181 of 295, top 62%highLMArena2026-10-08
Short-Story Creative Writing61.7%#38 of 39, top 98%highEpoch AI
Short-Story Creative Writing61.5%mediumEpoch AI
LMArena Multi-Turn1320#175 of 295, top 60%highLMArena2026-10-08
LiveBench Language50.7%#10 of 39, top 26%highEpoch AI
LiveBench Language38.3%lowEpoch AI
LiveBench Language46.3%mediumEpoch AI

API pricing by provider

o3-mini API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.10$4.40$0.552026-10-10
openai$1.10$4.40$0.552026-10-10
openrouter$1.10$4.40$0.552026-10-10

Compare o3-mini

Other OpenAI models

Frequently asked questions

How good is o3-mini?

o3-mini by OpenAI ranks 212th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.7. Its strongest category is instruction following, where it ranks 72nd. API pricing starts at $1.10 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.

How much does o3-mini cost?

o3-mini costs $1.10 per million input tokens and $4.40 per million output tokens on OpenAI's own API, with cached input at $0.55.

What is o3-mini's context window?

o3-mini accepts up to 200K tokens of input and can write up to 100K tokens in one response.

Is o3-mini open source?

No. o3-mini is proprietary and available only through OpenAI's API and partner platforms.

What are o3-mini's strengths and weaknesses?

Relative to other ranked models, o3-mini places best in instruction following, coding, knowledge and lowest in reasoning, long context, math.

What is o3-mini best at?

Its best category is instruction following, where it ranks 72nd on Noometry.