OpenAI, proprietary
GPT-4.1 mini
GPT-4.1 mini by OpenAI ranks 240th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.6. Its strongest category is agentic & tool use, where it ranks 55th. API pricing starts at $0.40 per million input tokens and $1.60 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #240 of 354
- Index score
- 33.6
- Evidence
- Confirmed 47 results
- Provider
- OpenAI
- Released
- April 14, 2025
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 1.05M
- Max output
- 33K
- Input price
- $0.40 / M
- Output price
- $1.60 / M
- Blended price
- $0.70 / M
- Output speed
- 86 tokens/s Kagi
- Value
- #97 of 219
- Knowledge cutoff
- April 2024
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 30.6
- Agentic & Tool Use 33.3
- Reasoning 10.8
- Math 24.1
- Knowledge 34.7
- Multimodal 35.8
- Multilingual 45.7
- Instruction Following 73.7
- Long Context 31.8
- Writing & Preference 48.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 30.6 | #293 | 7 |
| Agentic & Tool Use | 33.3 | #55 | 1 |
| Reasoning | 10.8 | #340 | 9 |
| Math | 24.1 | #270 | 5 |
| Knowledge | 34.7 | #194 | 5 |
| Multimodal | 35.8 | #82 | 1 |
| Multilingual | 45.7 | #166 | 1 |
| Instruction Following | 73.7 | #118 | 2 |
| Long Context | 31.8 | #275 | 2 |
| Writing & Preference | 48.6 | #199 | 5 |
Strengths and weaknesses
Categories where GPT-4.1 mini places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 33.3 | +2.9 | #55 of 154, top 36% |
| Instruction Following | 73.7 | +2.4 | #118 of 305, top 39% |
| Multilingual | 45.7 | −1.7 | #166 of 297, top 56% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 10.8 | −12.8 | #340 of 350, top 98% |
| Long Context | 31.8 | −9.1 | #275 of 296, top 93% |
| Coding | 30.6 | −8.1 | #293 of 340, top 87% |
Closest competitors
The models ranked just above and below GPT-4.1 mini. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen3.5-9B | #236 | 33.8 | $0.11 | — | Compare |
| Codellama 70b Instruct | #237 | 33.7 | — | — | Compare |
| Qwen3 8B | #238 | 33.7 | $0.31 | — | Compare |
| Grok-2 (Dec 2024) | #239 | 33.7 | — | — | Compare |
| GPT-5 Nano | #241 | 33.5 | $0.14 | 4 | Compare |
| Mercury 2.5 | #242 | 33.5 | $0.0675 | — | Compare |
| Mistral Small | #243 | 33.4 | $0.26 | 120 | Compare |
| Nova 2.0 Pro Preview | #244 | 33.4 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 23.9% | #34 of 39, top 88% | SWE-bench | 2025-07-20 | |
| Aider Polyglot | 32.4% | #32 of 44, top 73% | Epoch AI | ||
| SciCode | 40.4% | #76 of 121, top 63% | Epoch AI | ||
| WeirdML | 37.6% | #86 of 119, top 73% | Epoch AI | ||
| BigCodeBench Instruct | 48.9% | #6 of 64, top 10% | BigCodeBench | 2025-04-14 | |
| LMArena Coding | 1367 | #160 of 294, top 55% | LMArena | 2026-10-08 | |
| CadEval | 16% | #13 of 14, top 93% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 50.5% | #19 of 49, top 39% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0% | #75 of 83, top 91% | Epoch AI | ||
| Kagi LLM Benchmark | 48.6% | #68 of 99, top 69% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 3.5% | #81 of 83, top 98% | Epoch AI | ||
| CritPt | 0% | #110 of 134, top 83% | Epoch AI | ||
| Chess Puzzles | 7% | #87 of 129, top 68% | Epoch AI | 2026-07-15 | |
| LMArena Hard Prompts | 1349 | #161 of 297, top 55% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 7% | #66 of 74, top 90% | Epoch AI | 2026-08-27 | |
| DTBench | 68.8% | #94 of 151, top 63% | Epoch AI | ||
| LMCA | 21.1% | #92 of 125, top 74% | Epoch AI | ||
| Epoch Capabilities Index | 135.01 | #127 of 213, top 60% | Epoch AI | 2025-04-14 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 6.7% | #76 of 81, top 94% | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 44.7% | #115 of 173, top 67% | Epoch AI | 2025-04-14 | |
| Omni-MATH | 49.1% | #16 of 57, top 29% | HELM Capabilities | ||
| LMArena Math | 1343 | #166 of 285, top 59% | LMArena | 2026-10-08 | |
| MATH Level 5 | 87.3% | #18 of 79, top 23% | Epoch AI | 2025-04-14 | |
| FrontierMath (Feb 2025 set) | 4.5% | #48 of 68, top 71% | Epoch AI | 2025-04-14 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 65.8% | #104 of 186, top 56% | Epoch AI | 2025-04-14 | |
| SimpleQA Verified | 12.7% | #71 of 77, top 93% | Epoch AI | 2026-08-31 | |
| MMLU-Pro | 78.3% | #22 of 58, top 38% | HELM Capabilities | ||
| GPQA (HELM) | 61.4% | #21 of 57, top 37% | HELM Capabilities | ||
| LMArena Expert | 1338 | #161 of 273, top 59% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1181 | #86 of 122, top 71% | LMArena | 2026-10-09 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1318 | #166 of 297, top 56% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1329 | #178 of 285, top 63% | LMArena | 2026-10-08 | |
| LMArena French | 1358 | #142 of 223, top 64% | LMArena | 2026-10-08 | |
| LMArena German | 1351 | #129 of 231, top 56% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1290 | #126 of 211, top 60% | LMArena | 2026-10-08 | |
| LMArena Korean | 1298 | #131 of 213, top 62% | LMArena | 2026-10-08 | |
| LMArena Russian | 1324 | #165 of 283, top 59% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1319 | #155 of 226, top 69% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 90.4% | #9 of 57, top 16% | HELM Capabilities | ||
| LMArena Instruction Following | 1333 | #156 of 298, top 53% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 44.4% | #39 of 47, top 83% | Epoch AI | ||
| LMArena Longer Query | 1344 | #156 of 291, top 54% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1340 | #162 of 297, top 55% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1300 | #165 of 295, top 56% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1147 | #87 of 115, top 76% | EQ-Bench | ||
| WildBench | 83.8% | #17 of 57, top 30% | HELM Capabilities | ||
| LMArena Multi-Turn | 1354 | #153 of 295, top 52% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.40 | $1.60 | $0.10 | 2026-10-10 |
| openai | $0.40 | $1.60 | $0.10 | 2026-10-10 |
| openrouter | $0.40 | $1.60 | $0.10 | 2026-10-10 |
Compare GPT-4.1 mini
- GPT-4.1 mini vs GPT-4o mini
- GPT-4.1 mini vs Grok-2 (Dec 2024)
- GPT-4.1 mini vs GPT-5 Nano
- GPT-4.1 mini vs Qwen3 8B
- GPT-4.1 mini vs Mercury 2.5
- GPT-4.1 mini vs Codellama 70b Instruct
- GPT-4.1 mini vs Mistral Small
- GPT-4.1 mini vs Claude Fable 5.1
- GPT-4.1 mini vs Gemini 3.8 Flash
- GPT-4.1 mini vs Kimi K3
- GPT-4.1 mini vs Grok 4.6
- GPT-4.1 mini vs Qwen3.8 Max
- GPT-4.1 mini vs GLM-5.3
- GPT-4.1 mini vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-4.1 mini?
GPT-4.1 mini by OpenAI ranks 240th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.6. Its strongest category is agentic & tool use, where it ranks 55th. API pricing starts at $0.40 per million input tokens and $1.60 per million output tokens, with a 1.05M-token context window.
How much does GPT-4.1 mini cost?
GPT-4.1 mini costs $0.40 per million input tokens and $1.60 per million output tokens on OpenAI's own API, with cached input at $0.10.
What is GPT-4.1 mini's context window?
GPT-4.1 mini accepts up to 1.05M tokens of input and can write up to 33K tokens in one response.
Is GPT-4.1 mini open source?
No. GPT-4.1 mini is proprietary and available only through OpenAI's API and partner platforms.
How fast is GPT-4.1 mini?
GPT-4.1 mini generated about 86 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are GPT-4.1 mini's strengths and weaknesses?
Relative to other ranked models, GPT-4.1 mini places best in agentic & tool use, instruction following, multilingual and lowest in reasoning, long context, coding.
What is GPT-4.1 mini best at?
Its best category is agentic & tool use, where it ranks 55th on Noometry.