OpenAI, proprietary
o3
o3 by OpenAI ranks 61st of 354 ranked models on the Noometry Index as of October 2026, with a score of 47.5. Its strongest category is long context, where it ranks 6th. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 200K-token context window.
Last verified
Specifications
- Noometry rank
- #61 of 354
- Index score
- 47.5
- Evidence
- Confirmed 63 results
- Provider
- OpenAI
- Released
- April 16, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 200K
- Max output
- 100K
- Input price
- $2 / M
- Output price
- $8 / M
- Blended price
- $3.50 / M
- Output speed
- 3 tokens/s Kagi
- Value
- #168 of 219
- Knowledge cutoff
- May 2024
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 46.8
- Agentic & Tool Use 34.5
- Reasoning 32.0
- Math 50.2
- Knowledge 54.6
- Multimodal 41.4
- Multilingual 51.7
- Instruction Following 72.8
- Long Context 53.3
- Writing & Preference 63.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 46.8 | #64 | 7 |
| Agentic & Tool Use | 34.5 | #44 | 4 |
| Reasoning | 32.0 | #78 | 11 |
| Math | 50.2 | #58 | 5 |
| Knowledge | 54.6 | #52 | 7 |
| Multimodal | 41.4 | #36 | 3 |
| Multilingual | 51.7 | #105 | 1 |
| Instruction Following | 72.8 | #127 | 2 |
| Long Context | 53.3 | #6 | 3 |
| Writing & Preference | 63.5 | #64 | 6 |
Strengths and weaknesses
Categories where o3 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 53.3 | +12.3 | #6 of 296, top 3% |
| Knowledge | 54.6 | +17.3 | #52 of 314, top 17% |
| Math | 50.2 | +13.6 | #58 of 327, top 18% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 72.8 | +1.5 | #127 of 305, top 42% |
| Multilingual | 51.7 | +4.3 | #105 of 297, top 36% |
| Agentic & Tool Use | 34.5 | +4.1 | #44 of 154, top 29% |
Closest competitors
The models ranked just above and below o3. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Kimi K2.5 | #57 | 48.1 | $0.90 | 66 | Compare |
| Step 5 Preview | #58 | 47.9 | $1.43 | — | Compare |
| GLM-5.1 | #59 | 47.8 | $2.15 | — | Compare |
| Kimi K2.6 | #60 | 47.7 | $1.71 | — | Compare |
| Qwen3.6 Plus | #62 | 47.5 | $1.13 | — | Compare |
| Inkling-Small | #63 | 46.5 | $0.64 | — | Compare |
| GPT-5 Pro | #64 | 46.4 | $41.25 | 5 | Compare |
| Grok 4.20 Multi-Agent | #65 | 46.2 | $1.56 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 62.3% | #27 of 32, top 85% | medium | Epoch AI | 2026-02-12 |
| SWE-bench Verified (bash only) | 58.4% | #21 of 39, top 54% | SWE-bench | 2025-07-26 | |
| Aider Polyglot | 76.9% | Epoch AI | |||
| Aider Polyglot | 81.3% | #4 of 44, top 10% | high | Epoch AI | |
| Aider Polyglot | 76.9% | medium | Epoch AI | ||
| GSO | 8.8% | #18 of 31, top 59% | high | Epoch AI | |
| WeirdML | 52.4% | #44 of 119, top 37% | high | Epoch AI | |
| LMArena Coding | 1408 | #132 of 294, top 45% | LMArena | 2026-10-08 | |
| CadEval | 74% | Best of 14 | medium | Epoch AI | |
| ALE-Bench | 933.55 | #46 of 105, top 44% | high | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 63% | #7 of 49, top 15% | prompt | Berkeley Function Calling Leaderboard | |
| GDPval | 30.8% | #7 of 11, top 64% | medium | Epoch AI | |
| DeepResearch Bench | 45.2% | #16 of 24, top 67% | medium | Epoch AI | |
| OSWorld | 23% | #7 of 8, top 88% | medium | Epoch AI | |
| LMArena Search | 1144 | #25 of 32, top 79% | LMArena | 2026-08-24 | |
| METR Time Horizons | 63.6% | Epoch AI | |||
| METR Time Horizons | 65.4% | #14 of 32, top 44% | medium | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 6.5% | #51 of 83, top 62% | high | Epoch AI | |
| ARC-AGI-2 | 2% | low | Epoch AI | ||
| ARC-AGI-2 | 3% | medium | Epoch AI | ||
| SimpleBench | 53.1% | #38 of 77, top 50% | high | Epoch AI | |
| Kagi LLM Benchmark | 67.6% | #27 of 99, top 28% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 60.8% | #50 of 83, top 61% | high | Epoch AI | |
| ARC-AGI-1 | 41.5% | low | Epoch AI | ||
| ARC-AGI-1 | 53.8% | medium | Epoch AI | ||
| CritPt | 1.4% | #77 of 134, top 58% | high | Epoch AI | |
| Chess Puzzles | 34% | high | Epoch AI | 2026-08-07 | |
| Chess Puzzles | 27% | low | Epoch AI | 2026-07-15 | |
| Chess Puzzles | 38% | #26 of 129, top 21% | medium | Epoch AI | 2026-08-07 |
| EnigmaEval | 11.9% | high | Epoch AI | ||
| EnigmaEval | 13.1% | #10 of 38, top 27% | medium | Epoch AI | |
| LMArena Hard Prompts | 1402 | #124 of 297, top 42% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 29% | #29 of 74, top 40% | high | Epoch AI | 2026-08-27 |
| Mystery Game Puzzles | 23% | medium | Epoch AI | 2026-08-27 | |
| DTBench | 84.8% | #54 of 151, top 36% | high | Epoch AI | |
| LMCA | 39.7% | #46 of 125, top 37% | high | Epoch AI | |
| Epoch Capabilities Index | 146.86 | #67 of 213, top 32% | Epoch AI | 2025-04-16 | |
| ForecastBench | 62.5 | Best of 72 | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 33.3% | #60 of 81, top 75% | high | Epoch AI | 2026-08-27 |
| FrontierMath (Tiers 1-3) | 19.3% | low | Epoch AI | 2026-08-27 | |
| FrontierMath (Tiers 1-3) | 29.8% | medium | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 83.9% | high | Epoch AI | 2025-04-16 | |
| OTIS Mock AIME 2024-2025 | 60% | low | Epoch AI | 2026-07-15 | |
| OTIS Mock AIME 2024-2025 | 84.4% | #71 of 173, top 42% | medium | Epoch AI | 2026-08-07 |
| Omni-MATH | 71.4% | #4 of 57, top 8% | HELM Capabilities | ||
| LMArena Math | 1426 | #93 of 285, top 33% | LMArena | 2026-10-08 | |
| MATH Level 5 | 97.8% | #4 of 79, top 6% | high | Epoch AI | 2025-04-16 |
| FrontierMath (Feb 2025 set) | 18.7% | #32 of 68, top 48% | high | Epoch AI | 2025-11-16 |
| FrontierMath (Feb 2025 set) | 9.7% | low | Epoch AI | 2025-11-17 | |
| FrontierMath (Feb 2025 set) | 16.9% | medium | Epoch AI | 2025-11-16 | |
| FrontierMath Tier 4 (v1) | 2.1% | #42 of 55, top 77% | high | Epoch AI | 2025-07-01 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 81.8% | #78 of 186, top 42% | high | Epoch AI | 2025-04-16 |
| GPQA Diamond | 79.8% | low | Epoch AI | 2026-07-15 | |
| GPQA Diamond | 80.8% | medium | Epoch AI | 2026-08-07 | |
| Humanity's Last Exam | 20.3% | #18 of 41, top 44% | high | Epoch AI | |
| Humanity's Last Exam | 19.2% | medium | Epoch AI | ||
| SimpleQA Verified | 49.4% | #26 of 77, top 34% | high | Epoch AI | 2026-08-27 |
| MMLU-Pro | 85.9% | #6 of 58, top 11% | HELM Capabilities | ||
| Confabulations (lower is better) | 14.4% | #16 of 51, top 32% | high reasoning | Lech Mazur benchmarks | |
| GPQA (HELM) | 75.3% | #4 of 57, top 8% | HELM Capabilities | ||
| LMArena Expert | 1402 | #120 of 273, top 44% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1214 | #73 of 122, top 60% | LMArena | 2026-10-09 | |
| GeoBench | 60% | high | Epoch AI | ||
| GeoBench | 74% | #10 of 25, top 40% | medium | Epoch AI | |
| VPCT | 52% | #7 of 24, top 30% | medium | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1401 | #105 of 297, top 36% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1437 | #114 of 285, top 40% | LMArena | 2026-10-08 | |
| LMArena French | 1430 | #92 of 223, top 42% | LMArena | 2026-10-08 | |
| LMArena German | 1420 | #77 of 231, top 34% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1403 | #59 of 211, top 28% | LMArena | 2026-10-08 | |
| LMArena Korean | 1370 | #87 of 213, top 41% | LMArena | 2026-10-08 | |
| LMArena Russian | 1406 | #98 of 283, top 35% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1395 | #116 of 226, top 52% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 86.9% | #16 of 57, top 29% | HELM Capabilities | ||
| LMArena Instruction Following | 1368 | #131 of 298, top 44% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 88.9% | #6 of 47, top 13% | medium | Epoch AI | |
| CL-bench | 17.8% | #12 of 19, top 64% | high | Epoch AI | |
| LMArena Longer Query | 1372 | #139 of 291, top 48% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1410 | #110 of 297, top 38% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1359 | #122 of 295, top 42% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 83.9% | #5 of 39, top 13% | medium | Epoch AI | |
| EQ-Bench Creative Writing | 1676 | #35 of 115, top 31% | EQ-Bench | ||
| WildBench | 86.1% | #4 of 57, top 8% | HELM Capabilities | ||
| LMArena Multi-Turn | 1405 | #117 of 295, top 40% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $2 | $8 | $0.50 | 2026-10-10 |
| openai | $2 | $8 | $0.50 | 2026-10-10 |
| openrouter | $2 | $8 | $0.50 | 2026-10-10 |
Compare o3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is o3?
o3 by OpenAI ranks 61st of 354 ranked models on the Noometry Index as of October 2026, with a score of 47.5. Its strongest category is long context, where it ranks 6th. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 200K-token context window.
How much does o3 cost?
o3 costs $2 per million input tokens and $8 per million output tokens on OpenAI's own API, with cached input at $0.50.
What is o3's context window?
o3 accepts up to 200K tokens of input and can write up to 100K tokens in one response.
Is o3 open source?
No. o3 is proprietary and available only through OpenAI's API and partner platforms.
How fast is o3?
o3 generated about 3 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are o3's strengths and weaknesses?
Relative to other ranked models, o3 places best in long context, knowledge, math and lowest in instruction following, multilingual, agentic & tool use.
What is o3 best at?
Its best category is long context, where it ranks 6th on Noometry.