OpenAI, proprietary

# GPT-5.1-Codex

> GPT-5.1-Codex by OpenAI, released November 2025. Ranked #186 of 354 with a Noometry Index of 38.6. API: $1.25 in / $10 out per M tokens. 400K context. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/gpt-5-1-codex
- Last updated: 2026-10-10
- Title: GPT-5.1-Codex Benchmarks, Price & Rank (October 2026)

GPT-5.1-Codex by OpenAI ranks 186th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is agentic & tool use, where it ranks 33rd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #186 of 354
- **Index score:** 38.6
- **Evidence:** Reported 6 results
- **Provider:** [OpenAI](https://noometry.com/providers/openai)
- **Released:** November 12, 2025
- **Weights:** Proprietary
- **Reasoning:** Yes
- **Context window:** 400K
- **Max output:** 128K
- **Input price:** $1.25 / M
- **Output price:** $10 / M
- **Blended price:** $3.44 / M
- **Output speed:** Not measured
- **Value:** #178 of 219
- **Knowledge cutoff:** September 2024
- **Input:** text, image, audio

## Category scores

Each category score combines every public result we have in that category.

GPT-5.1-Codex category scores

1.  Coding 41.9
2.  Agentic & Tool Use 38.0
3.  Math 30.3
4.  2530354045

GPT-5.1-Codex category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 41.9 | #116 | 2 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 38.0 | #33 | 1 |
| [Math](https://noometry.com/best/math) | 30.3 | #235 | 1 |

## Strengths and weaknesses

Categories where GPT-5.1-Codex places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

GPT-5.1-Codex: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 38.0 | +7.6 | #33 of 154, top 22% |

### Weakest categories

GPT-5.1-Codex: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Math](https://noometry.com/best/math) | 30.3 | −6.2 | #235 of 327, top 72% |

## Closest competitors

The models ranked just above and below GPT-5.1-Codex. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5.1-Codex
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Qwen3.6 Flash](https://noometry.com/models/qwen3-6-flash) | #182 | 38.8 | $0.42 | — | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-qwen3-6-flash) |
| [Olmo 3 32b Think](https://noometry.com/models/olmo-3-32b-think) | #183 | 38.7 | — | — | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-olmo-3-32b-think) |
| [Hunyuan Large 2025 02 10](https://noometry.com/models/hunyuan-large) | #184 | 38.6 | — | — | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-hunyuan-large) |
| [Trinity Large Thinking](https://noometry.com/models/trinity-large-thinking) | #185 | 38.6 | $0.39 | — | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-trinity-large-thinking) |
| [Sonar](https://noometry.com/models/sonar) | #187 | 38.5 | $1 | — | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-sonar) |
| [MiniMax-M2.5](https://noometry.com/models/minimax-m2-5) | #188 | 38.3 | $0.52 | 49 | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-minimax-m2-5) |
| [Nova Premier 1.0](https://noometry.com/models/nova-premier-1-0) | #189 | 38.3 | $5 | 9 | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-nova-premier-1-0) |
| [Qwen3-Coder 480B-A35B Instruct](https://noometry.com/models/qwen3-coder-480b-a35b-instruct) | #190 | 38.1 | $3 | 67 | [Compare](https://noometry.com/compare/gpt-5-1-codex-vs-qwen3-coder-480b-a35b-instruct) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

GPT-5.1-Codex Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [SWE-bench Verified (bash only)](https://noometry.com/benchmarks/swe-bench-bash-only) | 66% | #15 of 39, top 39% | medium | [SWE-bench](https://www.swebench.com/) | 2025-11-24 |
| [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev) | 1337 | #92 of 113, top 82% |  | [LMArena](https://lmarena.ai/leaderboard/webdev) | 2026-10-08 |
| [ALE-Bench](https://noometry.com/benchmarks/ale-bench) | 1,209 |  |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [ALE-Bench](https://noometry.com/benchmarks/ale-bench) | 1,245 | #29 of 105, top 28% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Agentic & Tool Use

GPT-5.1-Codex Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Terminal-Bench](https://noometry.com/benchmarks/terminal-bench) | 57.8% |  |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Terminal-Bench](https://noometry.com/benchmarks/terminal-bench) | 60.4% | #13 of 41, top 32% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [METR Time Horizons](https://noometry.com/benchmarks/metr-time-horizons) | 70.8% | #9 of 32, top 29% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Math

GPT-5.1-Codex Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [ProofBench](https://noometry.com/benchmarks/proofbench) | 9% | #63 of 77, top 82% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## API pricing by provider

GPT-5.1-Codex API prices
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
| --- | --- | --- | --- | --- |
| [azure](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) | $1.25 | $10 | $0.13 | 2026-10-10 |
| [openrouter](https://openrouter.ai/openai/gpt-5.1-codex) | $1.25 | $10 | $0.13 | 2026-10-10 |

[All OpenAI API prices →](https://noometry.com/llm-pricing/openai) [Estimate your cost →](https://noometry.com/tools/cost-calculator)

## Compare GPT-5.1-Codex

-   [GPT-5.1-Codex vs GPT-5-Codex](https://noometry.com/compare/gpt-5-1-codex-vs-gpt-5-codex)
-   [GPT-5.1-Codex vs Trinity Large Thinking](https://noometry.com/compare/gpt-5-1-codex-vs-trinity-large-thinking)
-   [GPT-5.1-Codex vs Sonar](https://noometry.com/compare/gpt-5-1-codex-vs-sonar)
-   [GPT-5.1-Codex vs Hunyuan Large 2025 02 10](https://noometry.com/compare/gpt-5-1-codex-vs-hunyuan-large)
-   [GPT-5.1-Codex vs MiniMax-M2.5](https://noometry.com/compare/gpt-5-1-codex-vs-minimax-m2-5)
-   [GPT-5.1-Codex vs Olmo 3 32b Think](https://noometry.com/compare/gpt-5-1-codex-vs-olmo-3-32b-think)
-   [GPT-5.1-Codex vs Nova Premier 1.0](https://noometry.com/compare/gpt-5-1-codex-vs-nova-premier-1-0)
-   [GPT-5.1-Codex vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-gpt-5-1-codex)
-   [GPT-5.1-Codex vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-gpt-5-1-codex)
-   [GPT-5.1-Codex vs Kimi K3](https://noometry.com/compare/gpt-5-1-codex-vs-kimi-k3)
-   [GPT-5.1-Codex vs Grok 4.6](https://noometry.com/compare/gpt-5-1-codex-vs-grok-4-6)
-   [GPT-5.1-Codex vs Qwen3.8 Max](https://noometry.com/compare/gpt-5-1-codex-vs-qwen3-8-max)
-   [GPT-5.1-Codex vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-gpt-5-1-codex)
-   [GPT-5.1-Codex vs Muse Spark 1.3](https://noometry.com/compare/gpt-5-1-codex-vs-muse-spark-1-3)

## Other OpenAI models

-   [GPT-6 Astra](https://noometry.com/models/gpt-6-astra)70.8
-   [GPT-6.1 Sol](https://noometry.com/models/gpt-6-1-sol)65.6
-   [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol)65.0
-   [GPT-5.5 Pro](https://noometry.com/models/gpt-5-5-pro)64.3
-   [GPT-5.5](https://noometry.com/models/gpt-5-5)63.4
-   [GPT-6 Sol](https://noometry.com/models/gpt-6-sol)61.8
-   [GPT-5.4](https://noometry.com/models/gpt-5-4)59.4
-   [GPT-5.6 Terra](https://noometry.com/models/gpt-5-6-terra)59.2

## Frequently asked questions

### How good is GPT-5.1-Codex?

GPT-5.1-Codex by OpenAI ranks 186th of 354 ranked models on the Noometry Index as of October 2026, with a score of 38.6. Its strongest category is agentic & tool use, where it ranks 33rd. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 400K-token context window.

### How much does GPT-5.1-Codex cost?

GPT-5.1-Codex costs $1.25 per million input tokens and $10 per million output tokens on azure, with cached input at $0.13.

### What is GPT-5.1-Codex's context window?

GPT-5.1-Codex accepts up to 400K tokens of input and can write up to 128K tokens in one response.

### Is GPT-5.1-Codex open source?

No. GPT-5.1-Codex is proprietary and available only through OpenAI's API and partner platforms.

### What are GPT-5.1-Codex's strengths and weaknesses?

Relative to other ranked models, GPT-5.1-Codex places best in agentic & tool use and lowest in math.

### What is GPT-5.1-Codex best at?

Its best category is agentic & tool use, where it ranks 33rd on Noometry.

### Cite this page

Noometry. (2026). GPT-5.1-Codex benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/gpt-5-1-codex

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/gpt-5-1-codex.md).
