OpenAI, proprietary
GPT-4.1
GPT-4.1 by OpenAI ranks 219th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.9. Its strongest category is agentic & tool use, where it ranks 43rd. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #219 of 354
- Index score
- 35.9
- Evidence
- Confirmed 52 results
- Provider
- OpenAI
- Released
- April 14, 2025
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 1.05M
- Max output
- 33K
- Input price
- $2 / M
- Output price
- $8 / M
- Blended price
- $3.50 / M
- Output speed
- 116 tokens/s Kagi
- Value
- #184 of 219
- Knowledge cutoff
- April 2024
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 34.4
- Agentic & Tool Use 34.7
- Reasoning 11.7
- Math 22.3
- Knowledge 37.1
- Multimodal 38.2
- Multilingual 49.4
- Instruction Following 71.3
- Long Context 40.0
- Writing & Preference 57.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.4 | #238 | 6 |
| Agentic & Tool Use | 34.7 | #43 | 1 |
| Reasoning | 11.7 | #339 | 9 |
| Math | 22.3 | #280 | 5 |
| Knowledge | 37.1 | #160 | 7 |
| Multimodal | 38.2 | #67 | 2 |
| Multilingual | 49.4 | #133 | 1 |
| Instruction Following | 71.3 | #153 | 2 |
| Long Context | 40.0 | #163 | 2 |
| Writing & Preference | 57.6 | #125 | 5 |
Strengths and weaknesses
Categories where GPT-4.1 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 34.7 | +4.3 | #43 of 154, top 28% |
| Writing & Preference | 57.6 | +3.8 | #125 of 312, top 41% |
| Multilingual | 49.4 | +2.0 | #133 of 297, top 45% |
Closest competitors
The models ranked just above and below GPT-4.1. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Command A | #215 | 36.5 | $4.38 | 28 | Compare |
| Grok Build 0.1 | #216 | 36.4 | $1.25 | — | Compare |
| gpt-oss-120b | #217 | 36.3 | $0.0703 | 55 | Compare |
| Mistral Medium | #218 | 36.3 | $3 | 68 | Compare |
| Deepseek Coder v2 | #220 | 35.9 | — | — | Compare |
| C4ai Aya Expanse 32b | #221 | 35.9 | — | — | Compare |
| Llama 3.1 Nemotron 51b Instruct | #222 | 35.9 | — | — | Compare |
| Nemotron 4 340b Instruct | #223 | 35.9 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 48.5% | #31 of 32, top 97% | Epoch AI | 2026-02-08 | |
| SWE-bench Verified (bash only) | 39.6% | #30 of 39, top 77% | SWE-bench | 2025-07-26 | |
| Aider Polyglot | 52.4% | #21 of 44, top 48% | Epoch AI | ||
| WeirdML | 39% | #81 of 119, top 69% | Epoch AI | ||
| LMArena Coding | 1391 | #142 of 294, top 49% | LMArena | 2026-10-08 | |
| CadEval | 42% | #8 of 14, top 58% | Epoch AI | ||
| ALE-Bench | 558.1 | #84 of 105, top 80% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 54% | #15 of 49, top 31% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0.4% | #73 of 83, top 88% | Epoch AI | ||
| SimpleBench | 27% | #64 of 77, top 84% | Epoch AI | ||
| Kagi LLM Benchmark | 52.3% | #59 of 99, top 60% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 5.5% | #76 of 83, top 92% | Epoch AI | ||
| Chess Puzzles | 6% | #90 of 129, top 70% | Epoch AI | 2026-08-07 | |
| EnigmaEval | 2.2% | #30 of 38, top 79% | Epoch AI | ||
| LMArena Hard Prompts | 1384 | #135 of 297, top 46% | LMArena | 2026-10-08 | |
| DTBench | 68.3% | #96 of 151, top 64% | Epoch AI | ||
| LMCA | 25.6% | #85 of 125, top 68% | Epoch AI | ||
| Epoch Capabilities Index | 136.78 | #119 of 213, top 56% | Epoch AI | 2025-04-14 | |
| ForecastBench | 61.5 | #8 of 72, top 12% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 6% | #77 of 81, top 96% | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 38.3% | #116 of 173, top 68% | Epoch AI | 2025-04-14 | |
| Omni-MATH | 47.1% | #18 of 57, top 32% | HELM Capabilities | ||
| LMArena Math | 1370 | #150 of 285, top 53% | LMArena | 2026-10-08 | |
| MATH Level 5 | 83% | #23 of 79, top 30% | Epoch AI | 2025-04-14 | |
| FrontierMath (Feb 2025 set) | 5.5% | #45 of 68, top 67% | Epoch AI | 2025-04-14 | |
| FrontierMath Tier 4 (v1) | 0% | #50 of 55, top 91% | Epoch AI | 2025-07-01 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 66.9% | #103 of 186, top 56% | Epoch AI | 2025-04-14 | |
| Humanity's Last Exam | 5.4% | #35 of 41, top 86% | Epoch AI | ||
| SimpleQA Verified | 31.1% | #56 of 77, top 73% | Epoch AI | 2026-08-31 | |
| MMLU-Pro | 81.1% | #14 of 58, top 25% | HELM Capabilities | ||
| Vectara Hallucination Rate (lower is better) | 5.6% | #18 of 96, top 19% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 65.9% | #16 of 57, top 29% | HELM Capabilities | ||
| LMArena Expert | 1364 | #144 of 273, top 53% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1211 | #75 of 122, top 62% | LMArena | 2026-10-09 | |
| GeoBench | 72% | #11 of 25, top 44% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1370 | #133 of 297, top 45% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1382 | #148 of 285, top 52% | LMArena | 2026-10-08 | |
| LMArena French | 1382 | #130 of 223, top 59% | LMArena | 2026-10-08 | |
| LMArena German | 1381 | #110 of 231, top 48% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1319 | #113 of 211, top 54% | LMArena | 2026-10-08 | |
| LMArena Korean | 1339 | #110 of 213, top 52% | LMArena | 2026-10-08 | |
| LMArena Russian | 1377 | #130 of 283, top 46% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1376 | #128 of 226, top 57% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 83.8% | #25 of 57, top 44% | HELM Capabilities | ||
| LMArena Instruction Following | 1367 | #133 of 298, top 45% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 63.9% | #24 of 47, top 52% | Epoch AI | ||
| LMArena Longer Query | 1385 | #129 of 291, top 45% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1383 | #134 of 297, top 46% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1363 | #117 of 295, top 40% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1420 | #67 of 115, top 59% | EQ-Bench | ||
| WildBench | 85.4% | #10 of 57, top 18% | HELM Capabilities | ||
| LMArena Multi-Turn | 1398 | #121 of 295, top 42% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $2 | $8 | $0.50 | 2026-10-10 |
| openai | $2 | $8 | $0.50 | 2026-10-10 |
| openrouter | $2 | $8 | $0.50 | 2026-10-10 |
Compare GPT-4.1
- GPT-4.1 vs GPT-4.5
- GPT-4.1 vs Mistral Medium
- GPT-4.1 vs Deepseek Coder v2
- GPT-4.1 vs gpt-oss-120b
- GPT-4.1 vs C4ai Aya Expanse 32b
- GPT-4.1 vs Grok Build 0.1
- GPT-4.1 vs Llama 3.1 Nemotron 51b Instruct
- GPT-4.1 vs Claude Fable 5.1
- GPT-4.1 vs Gemini 3.8 Flash
- GPT-4.1 vs Kimi K3
- GPT-4.1 vs Grok 4.6
- GPT-4.1 vs Qwen3.8 Max
- GPT-4.1 vs GLM-5.3
- GPT-4.1 vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-4.1?
GPT-4.1 by OpenAI ranks 219th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.9. Its strongest category is agentic & tool use, where it ranks 43rd. API pricing starts at $2 per million input tokens and $8 per million output tokens, with a 1.05M-token context window.
How much does GPT-4.1 cost?
GPT-4.1 costs $2 per million input tokens and $8 per million output tokens on OpenAI's own API, with cached input at $0.50.
What is GPT-4.1's context window?
GPT-4.1 accepts up to 1.05M tokens of input and can write up to 33K tokens in one response.
Is GPT-4.1 open source?
No. GPT-4.1 is proprietary and available only through OpenAI's API and partner platforms.
How fast is GPT-4.1?
GPT-4.1 generated about 116 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are GPT-4.1's strengths and weaknesses?
Relative to other ranked models, GPT-4.1 places best in agentic & tool use, writing & preference, multilingual and lowest in reasoning, math, coding.
What is GPT-4.1 best at?
Its best category is agentic & tool use, where it ranks 43rd on Noometry.