OpenAI, proprietary
o3-mini
o3-mini by OpenAI ranks 212th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.7. Its strongest category is instruction following, where it ranks 72nd. API pricing starts at $1.10 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #212 of 354
- Index score
- 36.7
- Evidence
- Confirmed 51 results
- Provider
- OpenAI
- Released
- December 20, 2024
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 100K
- Input price
- $1.10 / M
- Output price
- $4.40 / M
- Blended price
- $1.93 / M
- Output speed
- Not measured
- Value
- #152 of 219
- Knowledge cutoff
- May 2024
- Input
- text
Category scores
Each category score combines every public result we have in that category.
- Coding 40.8
- Agentic & Tool Use 29.6
- Reasoning 16.3
- Math 28.1
- Knowledge 38.3
- Multilingual 45.7
- Instruction Following 75.1
- Long Context 33.8
- Writing & Preference 50.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 40.8 | #132 | 7 |
| Agentic & Tool Use | 29.6 | #84 | 1 |
| Reasoning | 16.3 | #305 | 11 |
| Math | 28.1 | #244 | 6 |
| Knowledge | 38.3 | #146 | 4 |
| Multilingual | 45.7 | #164 | 1 |
| Instruction Following | 75.1 | #72 | 2 |
| Long Context | 33.8 | #256 | 2 |
| Writing & Preference | 50.3 | #182 | 5 |
Strengths and weaknesses
Categories where o3-mini places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 75.1 | +3.8 | #72 of 305, top 24% |
| Coding | 40.8 | +2.1 | #132 of 340, top 39% |
| Knowledge | 38.3 | +1.0 | #146 of 314, top 47% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 16.3 | −7.3 | #305 of 350, top 88% |
| Long Context | 33.8 | −7.1 | #256 of 296, top 87% |
| Math | 28.1 | −8.5 | #244 of 327, top 75% |
Closest competitors
The models ranked just above and below o3-mini. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-4.5 | #208 | 37.2 | — | — | Compare |
| Yi-Lightning | #209 | 37.1 | — | — | Compare |
| Qwen Plus | #210 | 37.1 | $0.60 | 37 | Compare |
| Gemini 2.5 Flash-Lite | #211 | 37.0 | $0.18 | 172 | Compare |
| Llama 3.1 Nemotron Ultra 253b v1 | #213 | 36.7 | — | — | Compare |
| Granite 4.0 H Small | #214 | 36.5 | — | — | Compare |
| Command A | #215 | 36.5 | $4.38 | 28 | Compare |
| Grok Build 0.1 | #216 | 36.4 | $1.25 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 60.4% | #13 of 44, top 30% | high | Epoch AI | |
| Aider Polyglot | 53.8% | medium | Epoch AI | ||
| SciCode | 39.8% | #82 of 121, top 68% | high | Epoch AI | |
| GSO | 1.3% | #30 of 31, top 97% | high | Epoch AI | |
| GSO | 1.3% | low | Epoch AI | ||
| WeirdML | 43.7% | #64 of 119, top 54% | high | Epoch AI | |
| LiveBench Coding | 82.7% | #2 of 39, top 6% | high | Epoch AI | |
| LiveBench Coding | 61.5% | low | Epoch AI | ||
| LiveBench Coding | 65.4% | medium | Epoch AI | ||
| LMArena Coding | 1378 | #153 of 294, top 53% | high | LMArena | 2026-10-08 |
| CadEval | 54% | #6 of 14, top 43% | medium | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Cybench | 22.5% | #10 of 21, top 48% | medium | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 3% | #63 of 83, top 76% | high | Epoch AI | |
| ARC-AGI-2 | 0% | low | Epoch AI | ||
| ARC-AGI-2 | 2.1% | medium | Epoch AI | ||
| SimpleBench | 22.8% | #69 of 77, top 90% | high | Epoch AI | |
| ARC-AGI-1 | 34.5% | #63 of 83, top 76% | high | Epoch AI | |
| ARC-AGI-1 | 14.5% | low | Epoch AI | ||
| ARC-AGI-1 | 22.3% | medium | Epoch AI | ||
| CritPt | 0.3% | #96 of 134, top 72% | high | Epoch AI | |
| Chess Puzzles | 17% | #66 of 129, top 52% | high | Epoch AI | 2025-12-08 |
| Chess Puzzles | 6% | low | Epoch AI | 2026-07-15 | |
| Chess Puzzles | 9% | medium | Epoch AI | 2026-08-07 | |
| LiveBench Reasoning | 89.6% | #4 of 39, top 11% | high | Epoch AI | |
| LiveBench Reasoning | 69.8% | low | Epoch AI | ||
| LiveBench Reasoning | 86.3% | medium | Epoch AI | ||
| LMArena Hard Prompts | 1366 | #149 of 297, top 51% | high | LMArena | 2026-10-08 |
| Mystery Game Puzzles | 7% | #68 of 74, top 92% | high | Epoch AI | 2026-08-27 |
| DTBench | 68.8% | #95 of 151, top 63% | high | Epoch AI | |
| LiveBench Data Analysis | 70.6% | #4 of 39, top 11% | high | Epoch AI | |
| LiveBench Data Analysis | 62% | low | Epoch AI | ||
| LiveBench Data Analysis | 66.6% | medium | Epoch AI | ||
| LMCA | 19% | #94 of 125, top 76% | high | Epoch AI | |
| Epoch Capabilities Index | 140.34 | #107 of 213, top 51% | Epoch AI | 2025-01-31 | |
| ForecastBench | 59.6 | #37 of 72, top 52% | Epoch AI | ||
| LiveBench | 75.9% | #4 of 39, top 11% | high | Epoch AI | |
| LiveBench | 62.5% | low | Epoch AI | ||
| LiveBench | 70% | medium | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 18.6% | #72 of 81, top 89% | high | Epoch AI | 2026-06-11 |
| FrontierMath (Tiers 1-3) | 3.9% | low | Epoch AI | 2026-08-27 | |
| FrontierMath (Tiers 1-3) | 10.5% | medium | Epoch AI | 2026-08-27 | |
| FrontierMath Tier 4 | 0% | #63 of 63, top 100% | high | Epoch AI | 2026-06-11 |
| OTIS Mock AIME 2024-2025 | 76.9% | #84 of 173, top 49% | high | Epoch AI | 2025-02-27 |
| OTIS Mock AIME 2024-2025 | 44.4% | low | Epoch AI | 2026-07-20 | |
| OTIS Mock AIME 2024-2025 | 63.9% | medium | Epoch AI | 2025-02-25 | |
| LiveBench Math | 77.3% | #7 of 39, top 18% | high | Epoch AI | |
| LiveBench Math | 63.1% | low | Epoch AI | ||
| LiveBench Math | 72.4% | medium | Epoch AI | ||
| LMArena Math | 1396 | #132 of 285, top 47% | high | LMArena | 2026-10-08 |
| MATH Level 5 | 96.5% | #8 of 79, top 11% | high | Epoch AI | 2025-02-13 |
| MATH Level 5 | 95.2% | medium | Epoch AI | 2025-01-31 | |
| FrontierMath (Feb 2025 set) | 12.4% | #36 of 68, top 53% | high | Epoch AI | 2025-11-16 |
| FrontierMath (Feb 2025 set) | 8.1% | medium | Epoch AI | 2025-03-06 | |
| FrontierMath Tier 4 (v1) | 4.2% | #34 of 55, top 62% | high | Epoch AI | 2025-07-01 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 77% | #86 of 186, top 47% | high | Epoch AI | 2025-02-13 |
| GPQA Diamond | 68.2% | low | Epoch AI | 2026-07-20 | |
| GPQA Diamond | 74.3% | medium | Epoch AI | 2025-01-31 | |
| SimpleQA Verified | 15.3% | #69 of 77, top 90% | high | Epoch AI | 2026-08-31 |
| Confabulations (lower is better) | 17.9% | #28 of 51, top 55% | medium reasoning | Lech Mazur benchmarks | |
| LMArena Expert | 1364 | #145 of 273, top 54% | high | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1319 | #164 of 297, top 56% | high | LMArena | 2026-10-08 |
| LMArena Chinese | 1379 | #152 of 285, top 54% | high | LMArena | 2026-10-08 |
| LMArena French | 1334 | #150 of 223, top 68% | high | LMArena | 2026-10-08 |
| LMArena German | 1303 | #149 of 231, top 65% | high | LMArena | 2026-10-08 |
| LMArena Japanese | 1286 | #129 of 211, top 62% | high | LMArena | 2026-10-08 |
| LMArena Korean | 1314 | #121 of 213, top 57% | high | LMArena | 2026-10-08 |
| LMArena Russian | 1304 | #175 of 283, top 62% | high | LMArena | 2026-10-08 |
| LMArena Spanish | 1321 | #154 of 226, top 69% | high | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 84.4% | #3 of 39, top 8% | high | Epoch AI | |
| LiveBench Instruction Following | 80.1% | low | Epoch AI | ||
| LiveBench Instruction Following | 83.2% | medium | Epoch AI | ||
| LMArena Instruction Following | 1337 | #151 of 298, top 51% | high | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 50% | #35 of 47, top 75% | medium | Epoch AI | |
| LMArena Longer Query | 1343 | #157 of 291, top 54% | high | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1337 | #168 of 297, top 57% | high | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1286 | #181 of 295, top 62% | high | LMArena | 2026-10-08 |
| Short-Story Creative Writing | 61.7% | #38 of 39, top 98% | high | Epoch AI | |
| Short-Story Creative Writing | 61.5% | medium | Epoch AI | ||
| LMArena Multi-Turn | 1320 | #175 of 295, top 60% | high | LMArena | 2026-10-08 |
| LiveBench Language | 50.7% | #10 of 39, top 26% | high | Epoch AI | |
| LiveBench Language | 38.3% | low | Epoch AI | ||
| LiveBench Language | 46.3% | medium | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $1.10 | $4.40 | $0.55 | 2026-10-10 |
| openai | $1.10 | $4.40 | $0.55 | 2026-10-10 |
| openrouter | $1.10 | $4.40 | $0.55 | 2026-10-10 |
Compare o3-mini
- o3-mini vs o1-mini
- o3-mini vs Gemini 2.5 Flash-Lite
- o3-mini vs Llama 3.1 Nemotron Ultra 253b v1
- o3-mini vs Qwen Plus
- o3-mini vs Granite 4.0 H Small
- o3-mini vs Yi-Lightning
- o3-mini vs Command A
- o3-mini vs Claude Fable 5.1
- o3-mini vs Gemini 3.8 Flash
- o3-mini vs Kimi K3
- o3-mini vs Grok 4.6
- o3-mini vs Qwen3.8 Max
- o3-mini vs GLM-5.3
- o3-mini vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is o3-mini?
o3-mini by OpenAI ranks 212th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.7. Its strongest category is instruction following, where it ranks 72nd. API pricing starts at $1.10 per million input tokens and $4.40 per million output tokens, with a 200K-token context window.
How much does o3-mini cost?
o3-mini costs $1.10 per million input tokens and $4.40 per million output tokens on OpenAI's own API, with cached input at $0.55.
What is o3-mini's context window?
o3-mini accepts up to 200K tokens of input and can write up to 100K tokens in one response.
Is o3-mini open source?
No. o3-mini is proprietary and available only through OpenAI's API and partner platforms.
What are o3-mini's strengths and weaknesses?
Relative to other ranked models, o3-mini places best in instruction following, coding, knowledge and lowest in reasoning, long context, math.
What is o3-mini best at?
Its best category is instruction following, where it ranks 72nd on Noometry.