OpenAI, proprietary
GPT-5.2
GPT-5.2 by OpenAI ranks 34th of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.1. Its strongest category is multimodal, where it ranks 7th. API pricing starts at $1.75 per million input tokens and $14 per million output tokens, with a 400K-token context window.
Last verified
Specifications
- Noometry rank
- #34 of 354
- Index score
- 54.1
- Evidence
- Confirmed 67 results
- Provider
- OpenAI
- Released
- December 11, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 400K
- Max output
- 128K
- Input price
- $1.75 / M
- Output price
- $14 / M
- Blended price
- $4.81 / M
- Output speed
- 15 tokens/s Kagi
- Value
- #179 of 219
- Knowledge cutoff
- August 2025
- Input
- text, image
Category scores
Each category score combines every public result we have in that category.
- Coding 51.6
- Agentic & Tool Use 40.2
- Reasoning 50.2
- Math 60.0
- Knowledge 59.3
- Multimodal 51.3
- Multilingual 53.4
- Instruction Following 74.7
- Long Context 44.0
- Writing & Preference 66.8
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 51.6 | #37 | 7 |
| Agentic & Tool Use | 40.2 | #24 | 9 |
| Reasoning | 50.2 | #35 | 12 |
| Math | 60.0 | #38 | 6 |
| Knowledge | 59.3 | #32 | 5 |
| Multimodal | 51.3 | #7 | 3 |
| Multilingual | 53.4 | #67 | 1 |
| Instruction Following | 74.7 | #89 | 1 |
| Long Context | 44.0 | #78 | 2 |
| Writing & Preference | 66.8 | #32 | 4 |
Strengths and weaknesses
Categories where GPT-5.2 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 51.3 | +12.7 | #7 of 128, top 6% |
| Reasoning | 50.2 | +26.6 | #35 of 350, top 10% |
| Knowledge | 59.3 | +21.9 | #32 of 314, top 11% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 74.7 | +3.4 | #89 of 305, top 30% |
| Long Context | 44.0 | +3.1 | #78 of 296, top 27% |
| Multilingual | 53.4 | +6.0 | #67 of 297, top 23% |
Closest competitors
The models ranked just above and below GPT-5.2. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-5.6 Luna | #30 | 54.6 | $0.45 | 12 | Compare |
| DeepSeek V4 Pro | #31 | 54.3 | $0.99 | 16 | Compare |
| Gemini 3.5 Flash | #32 | 54.2 | $3.38 | — | Compare |
| Gemini 3.6 Flash | #33 | 54.1 | $1.50 | — | Compare |
| DeepSeek V4 Flash | #35 | 53.6 | $0.26 | 6 | Compare |
| GPT-6 Luna | #36 | 53.3 | $0.20 | — | Compare |
| Grok 4.7 | #37 | 53.1 | $3 | — | Compare |
| DeepSeek V4.1 Flash | #38 | 52.8 | $0.26 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 73.8% | #17 of 32, top 54% | high | Epoch AI | 2026-02-12 |
| SWE-bench Verified (bash only) | 72.8% | #7 of 39, top 18% | high | SWE-bench | 2026-02-17 |
| SWE-bench Verified (bash only) | 71.8% | high | SWE-bench | 2025-12-11 | |
| LMArena WebDev | 1416 | #66 of 113, top 59% | LMArena | 2026-10-08 | |
| SWE-bench Multilingual | 66.7% | #9 of 13, top 70% | high | SWE-bench | 2026-02-13 |
| GSO | 27.4% | #11 of 31, top 36% | high | Epoch AI | |
| WeirdML | 49.6% | low | Epoch AI | ||
| WeirdML | 63.4% | medium | Epoch AI | ||
| WeirdML | 49.6% | none | Epoch AI | ||
| WeirdML | 72.2% | #16 of 119, top 14% | xhigh | Epoch AI | |
| LMArena Coding | 1447 | #90 of 294, top 31% | LMArena | 2026-10-08 | |
| LMArena Coding | 1442 | high | LMArena | 2026-10-08 | |
| ALE-Bench | 1,294 | #27 of 105, top 26% | high | Epoch AI | |
| ALE-Bench | 1,250 | medium | Epoch AI | ||
| AlgoTune | 2.05 | Best of 18 | medium | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 64.9% | #9 of 41, top 22% | Epoch AI | ||
| Terminal-Bench | 64.9% | medium | Epoch AI | ||
| Berkeley Function Calling Leaderboard | 55.9% | #13 of 49, top 27% | fc | Berkeley Function Calling Leaderboard | |
| GDPval | 49.7% | Best of 11 | none | Epoch AI | |
| Remote Labor Index | 2.1% | Epoch AI | |||
| Remote Labor Index | 2.5% | #10 of 14, top 72% | medium | Epoch AI | |
| τ²-bench Airline | 83% | #2 of 7, top 29% | high | τ²-bench | 2026-02-26 |
| τ²-bench Banking | 32.2% | #13 of 26, top 50% | high | τ²-bench | 2026-02-26 |
| τ²-bench Retail | 81.6% | #2 of 7, top 29% | high | τ²-bench | 2026-02-26 |
| τ²-bench Telecom | 89.7% | #5 of 7, top 72% | high | τ²-bench | 2026-02-26 |
| DeepResearch Bench | 41.1% | #20 of 24, top 84% | low | Epoch AI | |
| LMArena Search | 1172 | LMArena | 2026-08-24 | ||
| LMArena Search | 1207 | #11 of 32, top 35% | LMArena | 2026-08-24 | |
| METR Time Horizons | 75.3% | #4 of 32, top 13% | high | Epoch AI | |
| Vending-Bench 2 | 3,591 | #40 of 60, top 67% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0.8% | Epoch AI | |||
| ARC-AGI-2 | 43.3% | high | Epoch AI | ||
| ARC-AGI-2 | 9.7% | low | Epoch AI | ||
| ARC-AGI-2 | 26.7% | medium | Epoch AI | ||
| ARC-AGI-2 | 52.9% | #33 of 83, top 40% | xhigh | Epoch AI | |
| SimpleBench | 45.8% | #47 of 77, top 62% | Epoch AI | ||
| SimpleBench | 45.8% | high | Epoch AI | ||
| Kagi LLM Benchmark | 73.3% | #17 of 99, top 18% | Kagi LLM Benchmark | ||
| NYT Connections (extended) | 83.6% | #31 of 91, top 35% | xhigh reasoning | Lech Mazur benchmarks | |
| ARC-AGI-1 | 12.3% | Epoch AI | |||
| ARC-AGI-1 | 78.7% | high | Epoch AI | ||
| ARC-AGI-1 | 55.7% | low | Epoch AI | ||
| ARC-AGI-1 | 72.7% | medium | Epoch AI | ||
| ARC-AGI-1 | 12.3% | none | Epoch AI | ||
| ARC-AGI-1 | 86.2% | #35 of 83, top 43% | xhigh | Epoch AI | |
| Chess Puzzles | 40% | high | Epoch AI | 2025-12-11 | |
| Chess Puzzles | 23% | low | Epoch AI | 2025-12-11 | |
| Chess Puzzles | 40% | medium | Epoch AI | 2025-12-11 | |
| Chess Puzzles | 4% | none | Epoch AI | 2026-07-13 | |
| Chess Puzzles | 49% | #11 of 129, top 9% | xhigh | Epoch AI | 2025-12-15 |
| EnigmaEval | 10.4% | #14 of 38, top 37% | Epoch AI | ||
| EBR-Bench | 23% | #13 of 24, top 55% | xhigh | Epoch AI | 2026-06-26 |
| LMArena Hard Prompts | 1445 | #72 of 297, top 25% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1428 | high | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 23% | #36 of 74, top 49% | high | Epoch AI | 2026-08-06 |
| Mystery Game Puzzles | 10% | low | Epoch AI | 2026-08-27 | |
| Mystery Game Puzzles | 22% | medium | Epoch AI | 2026-08-28 | |
| Mystery Game Puzzles | 14% | none | Epoch AI | 2026-08-27 | |
| DTBench | 90.9% | #33 of 151, top 22% | xhigh | Epoch AI | |
| LMCA | 43.9% | #38 of 125, top 31% | xhigh | Epoch AI | |
| Epoch Capabilities Index | 153.45 | #41 of 213, top 20% | Epoch AI | 2025-12-11 | |
| ForecastBench | 60.1 | #31 of 72, top 44% | Epoch AI |
Math
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 88.2% | high | Epoch AI | 2025-12-11 | |
| GPQA Diamond | 82.7% | low | Epoch AI | 2025-12-11 | |
| GPQA Diamond | 87.9% | medium | Epoch AI | 2025-12-11 | |
| GPQA Diamond | 73.2% | none | Epoch AI | 2026-07-13 | |
| GPQA Diamond | 91.4% | #26 of 186, top 14% | xhigh | Epoch AI | 2025-12-13 |
| Humanity's Last Exam | 27.8% | #12 of 41, top 30% | Epoch AI | ||
| SimpleQA Verified | 34.3% | high | Epoch AI | 2026-08-27 | |
| SimpleQA Verified | 32.8% | low | Epoch AI | 2026-08-27 | |
| SimpleQA Verified | 32.7% | medium | Epoch AI | 2026-08-27 | |
| SimpleQA Verified | 37.1% | #45 of 77, top 59% | xhigh | Epoch AI | 2026-08-27 |
| Vectara Hallucination Rate (lower is better) | 10.8% | Vectara Hallucination Leaderboard | |||
| Vectara Hallucination Rate (lower is better) | 8.4% | #39 of 96, top 41% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1438 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1445 | #74 of 273, top 28% | high | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1268 | #39 of 122, top 32% | LMArena | 2026-10-09 | |
| LMArena Vision | 1246 | high | LMArena | 2026-10-09 | |
| VPCT | 67% | high | Epoch AI | ||
| VPCT | 84% | #2 of 24, top 9% | xhigh | Epoch AI | |
| Furniture Assembly | 38.3% | #16 of 31, top 52% | xhigh | Epoch AI | 2026-09-10 |
| LMArena Document | 1405 | #36 of 38, top 95% | high | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1425 | #67 of 297, top 23% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1405 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1460 | #91 of 285, top 32% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1447 | high | LMArena | 2026-10-08 | |
| LMArena French | 1438 | LMArena | 2026-10-08 | ||
| LMArena French | 1455 | #61 of 223, top 28% | high | LMArena | 2026-10-08 |
| LMArena German | 1448 | #46 of 231, top 20% | LMArena | 2026-10-08 | |
| LMArena German | 1444 | high | LMArena | 2026-10-08 | |
| LMArena Japanese | 1418 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1420 | #41 of 211, top 20% | high | LMArena | 2026-10-08 |
| LMArena Korean | 1392 | #66 of 213, top 31% | LMArena | 2026-10-08 | |
| LMArena Korean | 1374 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1415 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1440 | #52 of 283, top 19% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1433 | #78 of 226, top 35% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1413 | high | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1417 | #78 of 298, top 27% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1409 | high | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CL-bench | 18.2% | #11 of 19, top 58% | Epoch AI | ||
| CL-bench | 18.1% | high | Epoch AI | ||
| LMArena Longer Query | 1428 | #80 of 291, top 28% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1413 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1439 | #65 of 297, top 22% | LMArena | 2026-10-08 | |
| LMArena Text | 1416 | high | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1401 | #77 of 295, top 27% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1376 | high | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1703 | #30 of 115, top 27% | EQ-Bench | ||
| LMArena Multi-Turn | 1458 | #40 of 295, top 14% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1419 | high | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $1.75 | $14 | $0.13 | 2026-10-10 |
| openai | $1.75 | $14 | $0.17 | 2026-10-10 |
| openrouter | $1.75 | $14 | $0.17 | 2026-10-10 |
Compare GPT-5.2
- GPT-5.2 vs GPT-5.1
- GPT-5.2 vs Gemini 3.6 Flash
- GPT-5.2 vs DeepSeek V4 Flash
- GPT-5.2 vs Gemini 3.5 Flash
- GPT-5.2 vs GPT-6 Luna
- GPT-5.2 vs DeepSeek V4 Pro
- GPT-5.2 vs Grok 4.7
- GPT-5.2 vs Claude Fable 5.1
- GPT-5.2 vs Gemini 3.8 Flash
- GPT-5.2 vs Kimi K3
- GPT-5.2 vs Grok 4.6
- GPT-5.2 vs Qwen3.8 Max
- GPT-5.2 vs GLM-5.3
- GPT-5.2 vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-5.2?
GPT-5.2 by OpenAI ranks 34th of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.1. Its strongest category is multimodal, where it ranks 7th. API pricing starts at $1.75 per million input tokens and $14 per million output tokens, with a 400K-token context window.
How much does GPT-5.2 cost?
GPT-5.2 costs $1.75 per million input tokens and $14 per million output tokens on OpenAI's own API, with cached input at $0.17.
What is GPT-5.2's context window?
GPT-5.2 accepts up to 400K tokens of input and can write up to 128K tokens in one response.
Is GPT-5.2 open source?
No. GPT-5.2 is proprietary and available only through OpenAI's API and partner platforms.
How fast is GPT-5.2?
GPT-5.2 generated about 15 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are GPT-5.2's strengths and weaknesses?
Relative to other ranked models, GPT-5.2 places best in multimodal, reasoning, knowledge and lowest in instruction following, long context, multilingual.
What is GPT-5.2 best at?
Its best category is multimodal, where it ranks 7th on Noometry.