OpenAI, proprietary
GPT-6 Astra
GPT-6 Astra by OpenAI ranks 1st of 354 ranked models on the Noometry Index as of October 2026, with a score of 70.8. Its strongest category is knowledge, where it ranks 1st. API pricing starts at $10 per million input tokens and $50 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #1 of 354
- Index score
- 70.8
- Evidence
- Confirmed 56 results
- Provider
- OpenAI
- Released
- September 3, 2026
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1.05M
- Max output
- 128K
- Input price
- $10 / M
- Output price
- $50 / M
- Blended price
- $20 / M
- Output speed
- Not measured
- Value
- #206 of 219
- Knowledge cutoff
- April 2026
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 73.7
- Agentic & Tool Use 52.9
- Reasoning 85.1
- Math 93.5
- Knowledge 75.3
- Multimodal 55.0
- Multilingual 53.7
- Instruction Following 76.3
- Long Context 44.5
- Writing & Preference 75.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 73.7 | #2 | 9 |
| Agentic & Tool Use | 52.9 | #3 | 4 |
| Reasoning | 85.1 | #1 | 10 |
| Math | 93.5 | #2 | 5 |
| Knowledge | 75.3 | #1 | 5 |
| Multimodal | 55.0 | #3 | 3 |
| Multilingual | 53.7 | #61 | 1 |
| Instruction Following | 76.3 | #44 | 1 |
| Long Context | 44.5 | #62 | 1 |
| Writing & Preference | 75.3 | #7 | 4 |
Strengths and weaknesses
Categories where GPT-6 Astra places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 44.5 | +3.6 | #62 of 296, top 21% |
| Multilingual | 53.7 | +6.3 | #61 of 297, top 21% |
| Instruction Following | 76.3 | +5.0 | #44 of 305, top 15% |
Closest competitors
The models ranked just above and below GPT-6 Astra. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude Fable 5.1 | #2 | 69.0 | $20 | — | Compare |
| Claude Opus 5.5 | #3 | 68.6 | $8 | — | Compare |
| Claude Opus 5 | #4 | 67.8 | $10 | — | Compare |
| Claude Fable 5 | #5 | 66.8 | $20 | 25 | Compare |
| GPT-6.1 Sol | #6 | 65.6 | $4 | — | Compare |
| GPT-5.6 Sol | #7 | 65.0 | $8 | 10 | Compare |
| GPT-5.5 Pro | #8 | 64.3 | $67.50 | — | Compare |
| GPT-5.5 | #9 | 63.4 | $11.25 | 25 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| DeepSWE | 73.2% | high | Epoch AI | ||
| DeepSWE | 67% | low | Epoch AI | ||
| DeepSWE | 73.2% | max | Epoch AI | ||
| DeepSWE | 72.8% | medium | Epoch AI | ||
| DeepSWE | 74.1% | #2 of 29, top 7% | xhigh | Epoch AI | |
| FrontierCode | 53.3% | #4 of 37, top 11% | max | Epoch AI | |
| LMArena WebDev | 1786 | #2 of 113, top 2% | LMArena | 2026-10-08 | |
| FrontierSWE | 65.5% | Best of 18 | max | Epoch AI | |
| SciCode | 55.4% | high | Epoch AI | ||
| SciCode | 54.1% | low | Epoch AI | ||
| SciCode | 56.5% | #19 of 121, top 16% | max | Epoch AI | |
| SciCode | 54.2% | medium | Epoch AI | ||
| SciCode | 53.5% | none | Epoch AI | ||
| SciCode | 55.7% | xhigh | Epoch AI | ||
| GSO | 79.4% | #2 of 31, top 7% | Epoch AI | ||
| WeirdML | 92.9% | high | Epoch AI | ||
| WeirdML | 93.3% | max | Epoch AI | ||
| WeirdML | 93.6% | Best of 119 | promax | Epoch AI | |
| LMArena Coding | 1487 | #35 of 294, top 12% | LMArena | 2026-10-08 | |
| MirrorCode | 46.7% | #4 of 9, top 45% | high | Epoch AI | 2026-08-30 |
| ALE-Bench | 2,951 | Best of 105 | max | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 64.7% | #8 of 49, top 17% | Epoch AI | ||
| Remote Labor Index | 20.8% | Best of 14 | Epoch AI | ||
| BALROG | 68.3% | Best of 35 | max | Epoch AI | |
| GDP.pdf | 34.2% | Best of 36 | max | Epoch AI | |
| Vending-Bench 2 | 15,515 | Best of 60 | Epoch AI |
Reasoning
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 93.7% | #2 of 81, top 3% | max | Epoch AI | 2026-08-30 |
| FrontierMath Tier 4 | 97.6% | #2 of 63, top 4% | high | Epoch AI | 2026-08-30 |
| FrontierMath Tier 4 | 87.8% | low | Epoch AI | 2026-08-30 | |
| FrontierMath Tier 4 | 97.6% | max | Epoch AI | 2026-08-30 | |
| FrontierMath Tier 4 | 97.6% | medium | Epoch AI | 2026-08-30 | |
| FrontierMath Tier 4 | 82.9% | none | Epoch AI | 2026-08-30 | |
| FrontierMath Tier 4 | 97.6% | xhigh | Epoch AI | 2026-08-30 | |
| OTIS Mock AIME 2024-2025 | 100% | #9 of 173, top 6% | max | Epoch AI | 2026-08-30 |
| ProofBench | 99% | #7 of 77, top 10% | Epoch AI | ||
| LMArena Math | 1465 | #46 of 285, top 17% | LMArena | 2026-10-08 | |
| FrontierMath Erdős | 2.9% | Best of 7 | max | Epoch AI | 2026-08-28 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 95.8% | Best of 186 | max | Epoch AI | 2026-08-30 |
| Humanity's Last Exam | 54.8% | Best of 41 | Epoch AI | ||
| SimpleQA Verified | 75.6% | Best of 77 | max | Epoch AI | 2026-08-30 |
| Vectara Hallucination Rate (lower is better) | 8.7% | #41 of 96, top 43% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1483 | #38 of 273, top 14% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1281 | #29 of 122, top 24% | LMArena | 2026-10-09 | |
| Blueprint-Bench 2 | 49.7% | #3 of 31, top 10% | Epoch AI | ||
| Furniture Assembly | 80% | #3 of 31, top 10% | max | Epoch AI | 2026-09-10 |
| LMArena Document | 1468 | #13 of 38, top 35% | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1430 | #61 of 297, top 21% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1484 | #61 of 285, top 22% | LMArena | 2026-10-08 | |
| LMArena French | 1456 | #56 of 223, top 26% | LMArena | 2026-10-08 | |
| LMArena German | 1440 | #56 of 231, top 25% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1379 | #79 of 211, top 38% | LMArena | 2026-10-08 | |
| LMArena Korean | 1426 | #32 of 213, top 16% | LMArena | 2026-10-08 | |
| LMArena Russian | 1436 | #57 of 283, top 21% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1407 | #103 of 226, top 46% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1450 | #42 of 298, top 15% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1456 | #44 of 291, top 16% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1441 | #61 of 297, top 21% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1418 | #55 of 295, top 19% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 2173 | Best of 115 | EQ-Bench | ||
| LMArena Multi-Turn | 1448 | #58 of 295, top 20% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $10 | $50 | $1 | 2026-10-10 |
| bedrock | $10 | $50 | $1 | 2026-10-10 |
| openai | $10 | $50 | $1 | 2026-10-10 |
| openrouter | $10 | $50 | $1 | 2026-10-10 |
Compare GPT-6 Astra
- GPT-6 Astra vs Claude Fable 5.1
- GPT-6 Astra vs Claude Opus 5.5
- GPT-6 Astra vs Claude Opus 5
- GPT-6 Astra vs Claude Fable 5
- GPT-6 Astra vs GPT-6.1 Sol
- GPT-6 Astra vs GPT-5.6 Sol
- GPT-6 Astra vs Gemini 3.8 Flash
- GPT-6 Astra vs Kimi K3
- GPT-6 Astra vs Grok 4.6
- GPT-6 Astra vs Qwen3.8 Max
- GPT-6 Astra vs GLM-5.3
- GPT-6 Astra vs Muse Spark 1.3
- GPT-6 Astra vs DeepSeek V4 Pro
- GPT-6 Astra vs MiMo-V2.6-Pro
Other OpenAI models
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
- GPT-5.4 Pro58.9
Frequently asked questions
How good is GPT-6 Astra?
GPT-6 Astra by OpenAI ranks 1st of 354 ranked models on the Noometry Index as of October 2026, with a score of 70.8. Its strongest category is knowledge, where it ranks 1st. API pricing starts at $10 per million input tokens and $50 per million output tokens, with a 1.05M-token context window.
How much does GPT-6 Astra cost?
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens on OpenAI's own API, with cached input at $1.
What is GPT-6 Astra's context window?
GPT-6 Astra accepts up to 1.05M tokens of input and can write up to 128K tokens in one response.
Is GPT-6 Astra open source?
No. GPT-6 Astra is proprietary and available only through OpenAI's API and partner platforms.
What are GPT-6 Astra's strengths and weaknesses?
Relative to other ranked models, GPT-6 Astra places best in reasoning, knowledge, coding and lowest in long context, multilingual, instruction following.
What is GPT-6 Astra best at?
Its best category is knowledge, where it ranks 1st on Noometry.