OpenAI, proprietary

GPT-4.1 mini

GPT-4.1 mini by OpenAI ranks 240th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.6. Its strongest category is agentic & tool use, where it ranks 55th. API pricing starts at $0.40 per million input tokens and $1.60 per million output tokens, with a 1.05M-token context window.

Last verified

Specifications

Noometry rank
#240 of 354
Index score
33.6
Evidence
Confirmed 47 results
Provider
OpenAI
Released
April 14, 2025
Weights
Proprietary
Reasoning
No
Context window
1.05M
Max output
33K
Input price
$0.40 / M
Output price
$1.60 / M
Blended price
$0.70 / M
Output speed
86 tokens/s Kagi
Value
#97 of 219
Knowledge cutoff
April 2024
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

GPT-4.1 mini category scores
  1. Coding 30.6
  2. Agentic & Tool Use 33.3
  3. Reasoning 10.8
  4. Math 24.1
  5. Knowledge 34.7
  6. Multimodal 35.8
  7. Multilingual 45.7
  8. Instruction Following 73.7
  9. Long Context 31.8
  10. Writing & Preference 48.6
GPT-4.1 mini category ranks
CategoryScoreRankResults
Coding30.6#2937
Agentic & Tool Use33.3#551
Reasoning10.8#3409
Math24.1#2705
Knowledge34.7#1945
Multimodal35.8#821
Multilingual45.7#1661
Instruction Following73.7#1182
Long Context31.8#2752
Writing & Preference48.6#1995

Strengths and weaknesses

Categories where GPT-4.1 mini places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-4.1 mini: strongest categories
CategoryScorevs medianRank
Agentic & Tool Use33.3+2.9#55 of 154, top 36%
Instruction Following73.7+2.4#118 of 305, top 39%
Multilingual45.7−1.7#166 of 297, top 56%

Weakest categories

GPT-4.1 mini: weakest categories
CategoryScorevs medianRank
Reasoning10.8−12.8#340 of 350, top 98%
Long Context31.8−9.1#275 of 296, top 93%
Coding30.6−8.1#293 of 340, top 87%

Closest competitors

The models ranked just above and below GPT-4.1 mini. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-4.1 mini
ModelRankScoreBlended $/MSpeed
Qwen3.5-9B#23633.8$0.11—Compare
Codellama 70b Instruct#23733.7——Compare
Qwen3 8B#23833.7$0.31—Compare
Grok-2 (Dec 2024)#23933.7——Compare
GPT-5 Nano#24133.5$0.144Compare
Mercury 2.5#24233.5$0.0675—Compare
Mistral Small#24333.4$0.26120Compare
Nova 2.0 Pro Preview#24433.4——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-4.1 mini Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)23.9%#34 of 39, top 88%SWE-bench2025-07-20
Aider Polyglot32.4%#32 of 44, top 73%Epoch AI
SciCode40.4%#76 of 121, top 63%Epoch AI
WeirdML37.6%#86 of 119, top 73%Epoch AI
BigCodeBench Instruct48.9%#6 of 64, top 10%BigCodeBench2025-04-14
LMArena Coding1367#160 of 294, top 55%LMArena2026-10-08
CadEval16%#13 of 14, top 93%Epoch AI

Agentic & Tool Use

GPT-4.1 mini Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard50.5%#19 of 49, top 39%fcBerkeley Function Calling Leaderboard

Reasoning

GPT-4.1 mini Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20%#75 of 83, top 91%Epoch AI
Kagi LLM Benchmark48.6%#68 of 99, top 69%Kagi LLM Benchmark
ARC-AGI-13.5%#81 of 83, top 98%Epoch AI
CritPt0%#110 of 134, top 83%Epoch AI
Chess Puzzles7%#87 of 129, top 68%Epoch AI2026-07-15
LMArena Hard Prompts1349#161 of 297, top 55%LMArena2026-10-08
Mystery Game Puzzles7%#66 of 74, top 90%Epoch AI2026-08-27
DTBench68.8%#94 of 151, top 63%Epoch AI
LMCA21.1%#92 of 125, top 74%Epoch AI
Epoch Capabilities Index135.01#127 of 213, top 60%Epoch AI2025-04-14

Math

GPT-4.1 mini Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)6.7%#76 of 81, top 94%Epoch AI2026-08-27
OTIS Mock AIME 2024-202544.7%#115 of 173, top 67%Epoch AI2025-04-14
Omni-MATH49.1%#16 of 57, top 29%HELM Capabilities
LMArena Math1343#166 of 285, top 59%LMArena2026-10-08
MATH Level 587.3%#18 of 79, top 23%Epoch AI2025-04-14
FrontierMath (Feb 2025 set)4.5%#48 of 68, top 71%Epoch AI2025-04-14

Knowledge

GPT-4.1 mini Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond65.8%#104 of 186, top 56%Epoch AI2025-04-14
SimpleQA Verified12.7%#71 of 77, top 93%Epoch AI2026-08-31
MMLU-Pro78.3%#22 of 58, top 38%HELM Capabilities
GPQA (HELM)61.4%#21 of 57, top 37%HELM Capabilities
LMArena Expert1338#161 of 273, top 59%LMArena2026-10-08

Multimodal

GPT-4.1 mini Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1181#86 of 122, top 71%LMArena2026-10-09

Multilingual

GPT-4.1 mini Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1318#166 of 297, top 56%LMArena2026-10-08
LMArena Chinese1329#178 of 285, top 63%LMArena2026-10-08
LMArena French1358#142 of 223, top 64%LMArena2026-10-08
LMArena German1351#129 of 231, top 56%LMArena2026-10-08
LMArena Japanese1290#126 of 211, top 60%LMArena2026-10-08
LMArena Korean1298#131 of 213, top 62%LMArena2026-10-08
LMArena Russian1324#165 of 283, top 59%LMArena2026-10-08
LMArena Spanish1319#155 of 226, top 69%LMArena2026-10-08

Instruction Following

GPT-4.1 mini Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval90.4%#9 of 57, top 16%HELM Capabilities
LMArena Instruction Following1333#156 of 298, top 53%LMArena2026-10-08

Long Context

GPT-4.1 mini Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench44.4%#39 of 47, top 83%Epoch AI
LMArena Longer Query1344#156 of 291, top 54%LMArena2026-10-08

Writing & Preference

GPT-4.1 mini Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1340#162 of 297, top 55%LMArena2026-10-08
LMArena Creative Writing1300#165 of 295, top 56%LMArena2026-10-08
EQ-Bench Creative Writing1147#87 of 115, top 76%EQ-Bench
WildBench83.8%#17 of 57, top 30%HELM Capabilities
LMArena Multi-Turn1354#153 of 295, top 52%LMArena2026-10-08

API pricing by provider

GPT-4.1 mini API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.40$1.60$0.102026-10-10
openai$0.40$1.60$0.102026-10-10
openrouter$0.40$1.60$0.102026-10-10

Compare GPT-4.1 mini

Other OpenAI models

Frequently asked questions

How good is GPT-4.1 mini?

GPT-4.1 mini by OpenAI ranks 240th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.6. Its strongest category is agentic & tool use, where it ranks 55th. API pricing starts at $0.40 per million input tokens and $1.60 per million output tokens, with a 1.05M-token context window.

How much does GPT-4.1 mini cost?

GPT-4.1 mini costs $0.40 per million input tokens and $1.60 per million output tokens on OpenAI's own API, with cached input at $0.10.

What is GPT-4.1 mini's context window?

GPT-4.1 mini accepts up to 1.05M tokens of input and can write up to 33K tokens in one response.

Is GPT-4.1 mini open source?

No. GPT-4.1 mini is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-4.1 mini?

GPT-4.1 mini generated about 86 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-4.1 mini's strengths and weaknesses?

Relative to other ranked models, GPT-4.1 mini places best in agentic & tool use, instruction following, multilingual and lowest in reasoning, long context, coding.

What is GPT-4.1 mini best at?

Its best category is agentic & tool use, where it ranks 55th on Noometry.