OpenAI, proprietary

# GPT-4 Turbo

> GPT-4 Turbo by OpenAI, released November 2023. Ranked #292 of 354 with a Noometry Index of 30.5. API: $10 in / $30 out per M tokens. 128K context. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/gpt-4-turbo
- Last updated: 2026-10-10
- Title: GPT-4 Turbo Benchmarks, Price & Rank (October 2026)

GPT-4 Turbo by OpenAI ranks 292nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.5. Its strongest category is multimodal, where it ranks 110th. API pricing starts at $10 per million input tokens and $30 per million output tokens, with a 128K-token context window.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #292 of 354
- **Index score:** 30.5
- **Evidence:** Confirmed 36 results
- **Provider:** [OpenAI](https://noometry.com/providers/openai)
- **Released:** November 6, 2023
- **Weights:** Proprietary
- **Reasoning:** No
- **Context window:** 128K
- **Max output:** 4K
- **Input price:** $10 / M
- **Output price:** $30 / M
- **Blended price:** $15 / M
- **Output speed:** Not measured
- **Value:** #209 of 219
- **Knowledge cutoff:** December 2023
- **Input:** text, image

## Category scores

Each category score combines every public result we have in that category.

GPT-4 Turbo category scores

1.  Coding 33.8
2.  Reasoning 15.3
3.  Math 9.0
4.  Knowledge 24.3
5.  Multimodal 30.6
6.  Multilingual 40.5
7.  Instruction Following 65.8
8.  Long Context 38.0
9.  Writing & Preference 47.7
10.  020406080

GPT-4 Turbo category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 33.8 | #249 | 4 |
| [Reasoning](https://noometry.com/best/reasoning) | 15.3 | #317 | 5 |
| [Math](https://noometry.com/best/math) | 9.0 | #322 | 4 |
| [Knowledge](https://noometry.com/best/knowledge) | 24.3 | #268 | 3 |
| [Multimodal](https://noometry.com/best/multimodal) | 30.6 | #110 | 1 |
| [Multilingual](https://noometry.com/best/multilingual) | 40.5 | #216 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 65.8 | #216 | 1 |
| [Long Context](https://noometry.com/best/long-context) | 38.0 | #206 | 1 |
| [Writing & Preference](https://noometry.com/best/writing) | 47.7 | #206 | 3 |

## Strengths and weaknesses

Categories where GPT-4 Turbo places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

GPT-4 Turbo: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Writing & Preference](https://noometry.com/best/writing) | 47.7 | −6.1 | #206 of 312, top 67% |
| [Long Context](https://noometry.com/best/long-context) | 38.0 | −2.9 | #206 of 296, top 70% |
| [Instruction Following](https://noometry.com/best/instruction-following) | 65.8 | −5.5 | #216 of 305, top 71% |

### Weakest categories

GPT-4 Turbo: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Math](https://noometry.com/best/math) | 9.0 | −27.6 | #322 of 327, top 99% |
| [Reasoning](https://noometry.com/best/reasoning) | 15.3 | −8.3 | #317 of 350, top 91% |
| [Multimodal](https://noometry.com/best/multimodal) | 30.6 | −7.9 | #110 of 128, top 86% |

## Closest competitors

The models ranked just above and below GPT-4 Turbo. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-4 Turbo
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Llama 3.1-405B](https://noometry.com/models/llama-3-1-405b) | #288 | 30.7 | — | 78 | [Compare](https://noometry.com/compare/gpt-4-turbo-vs-llama-3-1-405b) |
| [Yi-1.5-34B](https://noometry.com/models/yi-1-5-34b) | #289 | 30.6 | — | — | [Compare](https://noometry.com/compare/gpt-4-turbo-vs-yi-1-5-34b) |
| [Codestral](https://noometry.com/models/codestral) | #290 | 30.6 | $0.45 | 271 | [Compare](https://noometry.com/compare/codestral-vs-gpt-4-turbo) |
| [Llama-3.3-70B-Instruct](https://noometry.com/models/llama-3-3-70b-instruct) | #291 | 30.6 | $0.16 | — | [Compare](https://noometry.com/compare/gpt-4-turbo-vs-llama-3-3-70b-instruct) |
| [Qwen1.5-32B](https://noometry.com/models/qwen1-5-32b) | #293 | 30.5 | — | — | [Compare](https://noometry.com/compare/gpt-4-turbo-vs-qwen1-5-32b) |
| [Amazon Nova Micro](https://noometry.com/models/amazon-nova-micro) | #294 | 30.4 | $0.0613 | — | [Compare](https://noometry.com/compare/amazon-nova-micro-vs-gpt-4-turbo) |
| [Olmo 7b Instruct](https://noometry.com/models/olmo-7b-instruct) | #295 | 30.3 | — | — | [Compare](https://noometry.com/compare/gpt-4-turbo-vs-olmo-7b-instruct) |
| [Magistral Small](https://noometry.com/models/magistral-small) | #296 | 30.2 | $0.75 | 0 | [Compare](https://noometry.com/compare/gpt-4-turbo-vs-magistral-small) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

GPT-4 Turbo Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [WeirdML](https://noometry.com/benchmarks/weirdml) | 18% | #107 of 119, top 90% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [BigCodeBench Instruct](https://noometry.com/benchmarks/bigcodebench-instruct) | 48.2% | #9 of 64, top 15% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-09 |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 1268 | #218 of 294, top 75% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [BigCodeBench Complete](https://noometry.com/benchmarks/bigcodebench-complete) | 58.2% | #9 of 66, top 14% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-04-09 |
| [HumanEval+](https://noometry.com/benchmarks/humaneval-plus) | 86.6% | #6 of 45, top 14% | april 2024 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |
| [MBPP+](https://noometry.com/benchmarks/mbpp-plus) | 73.3% | #9 of 38, top 24% | nov 2023 | [EvalPlus](https://evalplus.github.io/leaderboard.html) |  |

### Agentic & Tool Use

GPT-4 Turbo Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [METR Time Horizons](https://noometry.com/benchmarks/metr-time-horizons) | 36.7% | #27 of 32, top 85% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [METR Time Horizons](https://noometry.com/benchmarks/metr-time-horizons) | 28.9% |  |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [METR Time Horizons](https://noometry.com/benchmarks/metr-time-horizons) | 35.2% |  |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

GPT-4 Turbo Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [SimpleBench](https://noometry.com/benchmarks/simplebench) | 25.1% | #66 of 77, top 86% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 6% | #91 of 129, top 71% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-07-15 |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 1251 | #221 of 297, top 75% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [DTBench](https://noometry.com/benchmarks/dtbench) | 61.6% | #112 of 151, top 75% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMCA](https://noometry.com/benchmarks/lmca) | 9.8% | #112 of 125, top 90% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 126.46 |  |  | [Epoch AI](https://epoch.ai/eci) | 2024-01-25 |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 127.25 | #149 of 213, top 70% |  | [Epoch AI](https://epoch.ai/eci) | 2024-04-09 |
| [ForecastBench](https://noometry.com/benchmarks/forecastbench) | 59.4 | #39 of 72, top 55% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Math

GPT-4 Turbo Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [FrontierMath (Tiers 1-3)](https://noometry.com/benchmarks/frontiermath) | 0.7% | #78 of 81, top 97% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-27 |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 6.7% | #142 of 173, top 83% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-02-27 |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 1272 | #200 of 285, top 71% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [MATH Level 5](https://noometry.com/benchmarks/math-level-5) | 35.4% |  |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-27 |
| [MATH Level 5](https://noometry.com/benchmarks/math-level-5) | 40% |  |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-27 |
| [MATH Level 5](https://noometry.com/benchmarks/math-level-5) | 46.7% | #48 of 79, top 61% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-02-27 |

### Knowledge

GPT-4 Turbo Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 46.6% | #139 of 186, top 75% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-02-27 |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 42.4% |  |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-27 |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 42.3% |  |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-27 |
| [Confabulations](https://noometry.com/benchmarks/confabulations) (lower is better) | 28.4% | #45 of 51, top 89% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/confabulations) |  |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 1223 | #211 of 273, top 78% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 81.3% | #14 of 81, top 18% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 79.6% |  |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Multimodal

GPT-4 Turbo Multimodal benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Vision](https://noometry.com/benchmarks/arena-vision) | 1090 | #109 of 122, top 90% |  | [LMArena](https://lmarena.ai/leaderboard/vision) | 2026-10-09 |

### Multilingual

GPT-4 Turbo Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 1245 | #216 of 297, top 73% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 1242 | #212 of 285, top 75% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 1276 | #173 of 223, top 78% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 1259 | #170 of 231, top 74% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 1194 | #163 of 211, top 78% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 1187 | #169 of 213, top 80% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 1259 | #205 of 283, top 73% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 1260 | #178 of 226, top 79% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

GPT-4 Turbo Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 1249 | #212 of 298, top 72% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Long Context

GPT-4 Turbo Long Context benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 1254 | #219 of 291, top 76% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

GPT-4 Turbo Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 1272 | #215 of 297, top 73% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 1269 | #192 of 295, top 66% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 1267 | #214 of 295, top 73% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

## API pricing by provider

GPT-4 Turbo API prices
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
| --- | --- | --- | --- | --- |
| [azure](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) | $10 | $30 | — | 2026-10-10 |
| [openai](https://platform.openai.com/docs/models) | $10 | $30 | — | 2026-10-10 |
| [openrouter](https://openrouter.ai/openai/gpt-4-turbo) | $10 | $30 | — | 2026-10-10 |

[All OpenAI API prices →](https://noometry.com/llm-pricing/openai) [Estimate your cost →](https://noometry.com/tools/cost-calculator)

## Compare GPT-4 Turbo

-   [GPT-4 Turbo vs GPT-3.5-turbo](https://noometry.com/compare/gpt-3-5-turbo-vs-gpt-4-turbo)
-   [GPT-4 Turbo vs Llama-3.3-70B-Instruct](https://noometry.com/compare/gpt-4-turbo-vs-llama-3-3-70b-instruct)
-   [GPT-4 Turbo vs Qwen1.5-32B](https://noometry.com/compare/gpt-4-turbo-vs-qwen1-5-32b)
-   [GPT-4 Turbo vs Codestral](https://noometry.com/compare/codestral-vs-gpt-4-turbo)
-   [GPT-4 Turbo vs Amazon Nova Micro](https://noometry.com/compare/amazon-nova-micro-vs-gpt-4-turbo)
-   [GPT-4 Turbo vs Yi-1.5-34B](https://noometry.com/compare/gpt-4-turbo-vs-yi-1-5-34b)
-   [GPT-4 Turbo vs Olmo 7b Instruct](https://noometry.com/compare/gpt-4-turbo-vs-olmo-7b-instruct)
-   [GPT-4 Turbo vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-gpt-4-turbo)
-   [GPT-4 Turbo vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-gpt-4-turbo)
-   [GPT-4 Turbo vs Kimi K3](https://noometry.com/compare/gpt-4-turbo-vs-kimi-k3)
-   [GPT-4 Turbo vs Grok 4.6](https://noometry.com/compare/gpt-4-turbo-vs-grok-4-6)
-   [GPT-4 Turbo vs Qwen3.8 Max](https://noometry.com/compare/gpt-4-turbo-vs-qwen3-8-max)
-   [GPT-4 Turbo vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-gpt-4-turbo)
-   [GPT-4 Turbo vs Muse Spark 1.3](https://noometry.com/compare/gpt-4-turbo-vs-muse-spark-1-3)

## Other OpenAI models

-   [GPT-6 Astra](https://noometry.com/models/gpt-6-astra)70.8
-   [GPT-6.1 Sol](https://noometry.com/models/gpt-6-1-sol)65.6
-   [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol)65.0
-   [GPT-5.5 Pro](https://noometry.com/models/gpt-5-5-pro)64.3
-   [GPT-5.5](https://noometry.com/models/gpt-5-5)63.4
-   [GPT-6 Sol](https://noometry.com/models/gpt-6-sol)61.8
-   [GPT-5.4](https://noometry.com/models/gpt-5-4)59.4
-   [GPT-5.6 Terra](https://noometry.com/models/gpt-5-6-terra)59.2

## Frequently asked questions

### How good is GPT-4 Turbo?

GPT-4 Turbo by OpenAI ranks 292nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.5. Its strongest category is multimodal, where it ranks 110th. API pricing starts at $10 per million input tokens and $30 per million output tokens, with a 128K-token context window.

### How much does GPT-4 Turbo cost?

GPT-4 Turbo costs $10 per million input tokens and $30 per million output tokens on OpenAI's own API.

### What is GPT-4 Turbo's context window?

GPT-4 Turbo accepts up to 128K tokens of input and can write up to 4K tokens in one response.

### Is GPT-4 Turbo open source?

No. GPT-4 Turbo is proprietary and available only through OpenAI's API and partner platforms.

### What are GPT-4 Turbo's strengths and weaknesses?

Relative to other ranked models, GPT-4 Turbo places best in writing & preference, long context, instruction following and lowest in math, reasoning, multimodal.

### What is GPT-4 Turbo best at?

Its best category is multimodal, where it ranks 110th on Noometry.

### Cite this page

Noometry. (2026). GPT-4 Turbo benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/gpt-4-turbo

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/gpt-4-turbo.md).
