OpenAI, proprietary

GPT-4o

GPT-4o by OpenAI ranks 324th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.6. Its strongest category is multimodal, where it ranks 91st. API pricing starts at $2.50 per million input tokens and $10 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#324 of 354
Index score
28.6
Evidence
Confirmed 72 results
Provider
OpenAI
Released
May 13, 2024
Weights
Proprietary
Reasoning
No
Context window
128K
Max output
16K
Input price
$2.50 / M
Output price
$10 / M
Blended price
$4.38 / M
Output speed
Not measured
Value
#200 of 219
Knowledge cutoff
September 2023
Input
text, image

Category scores

Each category score combines every public result we have in that category.

GPT-4o category scores
  1. Coding 24.8
  2. Agentic & Tool Use 21.0
  3. Reasoning 9.4
  4. Math 10.6
  5. Knowledge 28.8
  6. Multimodal 34.5
  7. Multilingual 43.2
  8. Instruction Following 66.6
  9. Long Context 39.4
  10. Writing & Preference 52.6
GPT-4o category ranks
CategoryScoreRankResults
Coding24.8#32810
Agentic & Tool Use21.0#1414
Reasoning9.4#34311
Math10.6#3126
Knowledge28.8#2428
Multimodal34.5#914
Multilingual43.2#1861
Instruction Following66.6#2073
Long Context39.4#1792
Writing & Preference52.6#1666

Strengths and weaknesses

Categories where GPT-4o places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-4o: strongest categories
CategoryScorevs medianRank
Writing & Preference52.6−1.2#166 of 312, top 54%
Long Context39.4−1.5#179 of 296, top 61%
Multilingual43.2−4.2#186 of 297, top 63%

Weakest categories

GPT-4o: weakest categories
CategoryScorevs medianRank
Reasoning9.4−14.2#343 of 350, top 98%
Coding24.8−13.9#328 of 340, top 97%
Math10.6−25.9#312 of 327, top 96%

Closest competitors

The models ranked just above and below GPT-4o. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-4o
ModelRankScoreBlended $/MSpeed
Qwen2.5 7B Instruct#32029.0$0.31—Compare
Llama 3.2 3B#32128.9$0.12—Compare
Qwen1.5 4b Chat#32228.8——Compare
Llama 3-70B#32328.8—104Compare
Ministral 8B#32528.2$0.15—Compare
Gemma 3 4B#32628.1$0.0572Compare
GPT-4.1 nano#32727.9$0.18135Compare
Phi 3 Mini 4k Instruct#32827.9——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-4o Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified31%#32 of 32, top 100%Epoch AI2026-02-11
SWE-bench Verified (bash only)21.6%#35 of 39, top 90%SWE-bench2025-07-20
Aider Polyglot45.3%#24 of 44, top 55%Epoch AI
Aider Polyglot27.1%Epoch AI
Aider Polyglot18.2%Epoch AI
Aider Polyglot23.1%Epoch AI
GSO0%#31 of 31, top 100%Epoch AI
WeirdML25.1%#99 of 119, top 84%Epoch AI
BigCodeBench Instruct51.1%Best of 64BigCodeBench2024-05-13
BigCodeBench Instruct48%BigCodeBench2024-11-20
LiveBench Coding51.4%#16 of 39, top 42%Epoch AI
LiveBench Coding46.1%Epoch AI
LMArena Coding1297#199 of 294, top 68%LMArena2026-10-08
LMArena Coding1283LMArena2026-10-08
BigCodeBench Complete58.9%BigCodeBench2024-11-20
BigCodeBench Complete61.1%#3 of 66, top 5%BigCodeBench2024-05-13
CadEval26%#12 of 14, top 86%Epoch AI
HumanEval+87.2%#3 of 45, top 7%aug 2024EvalPlus
MBPP+72.2%#11 of 38, top 29%aug 2024EvalPlus

Agentic & Tool Use

GPT-4o Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
GDPval9.9%#11 of 11, top 100%Epoch AI
TheAgentCompany8.6%#8 of 14, top 58%Epoch AI
Cybench12.5%#14 of 21, top 67%Epoch AI
BALROG32.3%#16 of 35, top 46%Epoch AI
LMArena Search1006#32 of 32, top 100%LMArena2026-08-24
METR Time Horizons40.8%#26 of 32, top 82%Epoch AI

Reasoning

GPT-4o Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20%#77 of 83, top 93%Epoch AI
SimpleBench17.8%#75 of 77, top 98%Epoch AI
ARC-AGI-14.5%#79 of 83, top 96%Epoch AI
CritPt0%#112 of 134, top 84%Epoch AI
Chess Puzzles13%#73 of 129, top 57%Epoch AI2026-07-15
EnigmaEval0.8%#35 of 38, top 93%Epoch AI
LiveBench Reasoning53.9%Epoch AI
LiveBench Reasoning55.8%#15 of 39, top 39%Epoch AI
LMArena Hard Prompts1281#199 of 297, top 68%LMArena2026-10-08
LMArena Hard Prompts1264LMArena2026-10-08
DTBench64.5%#103 of 151, top 69%Epoch AI
LiveBench Data Analysis60.9%#14 of 39, top 36%Epoch AI
LiveBench Data Analysis56.1%Epoch AI
LMCA16.6%#102 of 125, top 82%Epoch AI
Epoch Capabilities Index128.97#143 of 213, top 68%Epoch AI2024-05-13
Epoch Capabilities Index128.76Epoch AI2024-08-06
Epoch Capabilities Index128.81Epoch AI2024-11-20
ForecastBench57.7#54 of 72, top 75%Epoch AI
LiveBench52.2%Epoch AI
LiveBench55.3%#15 of 39, top 39%Epoch AI

Math

GPT-4o Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)0.4%#80 of 81, top 99%Epoch AI2026-08-27
OTIS Mock AIME 2024-20256.3%Epoch AI2025-02-25
OTIS Mock AIME 2024-20256.3%Epoch AI2025-02-25
OTIS Mock AIME 2024-20256.4%#144 of 173, top 84%Epoch AI2025-02-25
Omni-MATH29.3%#40 of 57, top 71%HELM Capabilities
LiveBench Math42.9%Epoch AI
LiveBench Math49.5%#20 of 39, top 52%Epoch AI
LMArena Math1284LMArena2026-10-08
LMArena Math1285#189 of 285, top 67%LMArena2026-10-08
MATH Level 549.8%Epoch AI2025-02-05
MATH Level 553.3%#43 of 79, top 55%Epoch AI2025-01-27
MATH Level 551%Epoch AI2025-01-27
FrontierMath (Feb 2025 set)0.3%#65 of 68, top 96%Epoch AI2025-03-07
FrontierMath (Feb 2025 set)0.3%Epoch AI2025-03-06

Knowledge

GPT-4o Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond48.9%Epoch AI2025-01-27
GPQA Diamond47.9%Epoch AI2025-02-05
GPQA Diamond49.2%#128 of 186, top 69%Epoch AI2025-01-27
Humanity's Last Exam2.7%#41 of 41, top 100%Epoch AI
SimpleQA Verified26%#60 of 77, top 78%Epoch AI2026-08-31
MMLU-Pro71.3%#35 of 58, top 61%HELM Capabilities
Confabulations (lower is better)15.3%#18 of 51, top 36%Lech Mazur benchmarks
Confabulations (lower is better)17.2%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)9.6%#51 of 96, top 54%Vectara Hallucination Leaderboard
GPQA (HELM)52%#31 of 57, top 55%HELM Capabilities
LMArena Expert1241LMArena2026-10-08
LMArena Expert1250#196 of 273, top 72%LMArena2026-10-08
MMLU88.1%Best of 81Epoch AI
MMLU84.2%Epoch AI
MMLU84.3%Epoch AI

Multimodal

GPT-4o Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1137#102 of 122, top 84%LMArena2026-10-09
LMArena Vision1065LMArena2026-10-09
Video-MME71.9%#6 of 15, top 40%Epoch AI
Video-MME71.9%#6 of 15, top 40%Epoch AI
GeoBench71%#12 of 25, top 48%Epoch AI
VPCT40%#13 of 24, top 55%Epoch AI
ScienceQA88.5%Best of 6Epoch AI

Multilingual

GPT-4o Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1262LMArena2026-10-08
LMArena Non-English1283#186 of 297, top 63%LMArena2026-10-08
LMArena Chinese1254LMArena2026-10-08
LMArena Chinese1277#193 of 285, top 68%LMArena2026-10-08
LMArena French1304#160 of 223, top 72%LMArena2026-10-08
LMArena French1263LMArena2026-10-08
LMArena German1257LMArena2026-10-08
LMArena German1282#157 of 231, top 68%LMArena2026-10-08
LMArena Japanese1234LMArena2026-10-08
LMArena Japanese1257#138 of 211, top 66%LMArena2026-10-08
LMArena Korean1218LMArena2026-10-08
LMArena Korean1234#150 of 213, top 71%LMArena2026-10-08
LMArena Russian1272LMArena2026-10-08
LMArena Russian1286#186 of 283, top 66%LMArena2026-10-08
LMArena Spanish1292#163 of 226, top 73%LMArena2026-10-08
LMArena Spanish1269LMArena2026-10-08

Instruction Following

GPT-4o Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following64.9%Epoch AI
LiveBench Instruction Following68.6%#20 of 39, top 52%Epoch AI
IFEval81.7%#35 of 57, top 62%HELM Capabilities
LMArena Instruction Following1278#190 of 298, top 64%LMArena2026-10-08
LMArena Instruction Following1267LMArena2026-10-08

Long Context

GPT-4o Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench66.7%#19 of 47, top 41%Epoch AI
LMArena Longer Query1289#198 of 291, top 69%LMArena2026-10-08
LMArena Longer Query1283LMArena2026-10-08

Writing & Preference

GPT-4o Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1300#190 of 297, top 64%LMArena2026-10-08
LMArena Text1283LMArena2026-10-08
LMArena Creative Writing1275LMArena2026-10-08
LMArena Creative Writing1292#173 of 295, top 59%LMArena2026-10-08
Short-Story Creative Writing81.8%#11 of 39, top 29%Epoch AI
WildBench82.8%#19 of 57, top 34%HELM Capabilities
LMArena Multi-Turn1279LMArena2026-10-08
LMArena Multi-Turn1302#186 of 295, top 64%LMArena2026-10-08
LiveBench Language47.4%Epoch AI
LiveBench Language47.6%#14 of 39, top 36%Epoch AI

API pricing by provider

GPT-4o API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$2.50$10$1.252026-10-10
openai$2.50$10$1.252026-10-10
openrouter$2.50$10$1.252026-10-10

Compare GPT-4o

Other OpenAI models

Frequently asked questions

How good is GPT-4o?

GPT-4o by OpenAI ranks 324th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.6. Its strongest category is multimodal, where it ranks 91st. API pricing starts at $2.50 per million input tokens and $10 per million output tokens, with a 128K-token context window.

How much does GPT-4o cost?

GPT-4o costs $2.50 per million input tokens and $10 per million output tokens on OpenAI's own API, with cached input at $1.25.

What is GPT-4o's context window?

GPT-4o accepts up to 128K tokens of input and can write up to 16K tokens in one response.

Is GPT-4o open source?

No. GPT-4o is proprietary and available only through OpenAI's API and partner platforms.

What are GPT-4o's strengths and weaknesses?

Relative to other ranked models, GPT-4o places best in writing & preference, long context, multilingual and lowest in reasoning, coding, math.

What is GPT-4o best at?

Its best category is multimodal, where it ranks 91st on Noometry.