OpenAI, proprietary
GPT-6 Luna
GPT-6 Luna by OpenAI ranks 36th of 354 ranked models on the Noometry Index as of October 2026, with a score of 53.3. Its strongest category is math, where it ranks 15th. API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #36 of 354
- Index score
- 53.3
- Evidence
- Confirmed 42 results
- Provider
- OpenAI
- Released
- September 22, 2026
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1.05M
- Max output
- 128K
- Input price
- $0.10 / M
- Output price
- $0.50 / M
- Blended price
- $0.20 / M
- Output speed
- Not measured
- Value
- #26 of 219
- Knowledge cutoff
- May 2026
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 55.5
- Agentic & Tool Use 33.3
- Reasoning 48.2
- Math 76.1
- Knowledge 57.0
- Multimodal 42.4
- Multilingual 50.5
- Instruction Following 74.3
- Long Context 43.0
- Writing & Preference 58.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 55.5 | #25 | 5 |
| Agentic & Tool Use | 33.3 | #54 | 2 |
| Reasoning | 48.2 | #41 | 9 |
| Math | 76.1 | #15 | 5 |
| Knowledge | 57.0 | #41 | 3 |
| Multimodal | 42.4 | #30 | 3 |
| Multilingual | 50.5 | #117 | 1 |
| Instruction Following | 74.3 | #99 | 1 |
| Long Context | 43.0 | #111 | 1 |
| Writing & Preference | 58.3 | #119 | 3 |
Strengths and weaknesses
Categories where GPT-6 Luna places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 50.5 | +3.1 | #117 of 297, top 40% |
| Writing & Preference | 58.3 | +4.5 | #119 of 312, top 39% |
| Long Context | 43.0 | +2.1 | #111 of 296, top 38% |
Closest competitors
The models ranked just above and below GPT-6 Luna. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Gemini 3.5 Flash | #32 | 54.2 | $3.38 | — | Compare |
| Gemini 3.6 Flash | #33 | 54.1 | $1.50 | — | Compare |
| GPT-5.2 | #34 | 54.1 | $4.81 | 15 | Compare |
| DeepSeek V4 Flash | #35 | 53.6 | $0.26 | 6 | Compare |
| Grok 4.7 | #37 | 53.1 | $3 | — | Compare |
| DeepSeek V4.1 Flash | #38 | 52.8 | $0.26 | — | Compare |
| GPT-5.2 Pro | #39 | 52.3 | $57.75 | — | Compare |
| Gemini 3 Flash Preview | #40 | 52.3 | $1.13 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| DeepSWE | 59.3% | high | Model card (self-reported) | 2026-09-22 | |
| DeepSWE | 2.4% | low | Model card (self-reported) | 2026-09-22 | |
| DeepSWE | 66.6% | #14 of 29, top 49% | max | Model card (self-reported) | 2026-09-22 |
| DeepSWE | 44.5% | medium | Model card (self-reported) | 2026-09-22 | |
| DeepSWE | 61.3% | xhigh | Model card (self-reported) | 2026-09-22 | |
| FrontierCode | 42.4% | #18 of 37, top 49% | max | Epoch AI | |
| FrontierCode | 37.3% | high | Model card (self-reported) | 2026-09-22 | |
| FrontierCode | 25.7% | low | Model card (self-reported) | 2026-09-22 | |
| FrontierCode | 42.4% | max | Model card (self-reported) | 2026-09-22 | |
| FrontierCode | 35.5% | medium | Model card (self-reported) | 2026-09-22 | |
| FrontierCode | 37.1% | xhigh | Model card (self-reported) | 2026-09-22 | |
| LMArena WebDev | 1581 | #29 of 113, top 26% | LMArena | 2026-10-08 | |
| SciCode | 50.3% | high | Epoch AI | ||
| SciCode | 46.9% | low | Epoch AI | ||
| SciCode | 54.6% | #26 of 121, top 22% | max | Epoch AI | |
| SciCode | 50.9% | medium | Epoch AI | ||
| SciCode | 43.1% | none | Epoch AI | ||
| SciCode | 51.7% | xhigh | Epoch AI | ||
| LMArena Coding | 1439 | #98 of 294, top 34% | LMArena | 2026-10-08 | |
| ALE-Bench | 1,577 | #14 of 105, top 14% | xhigh | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 44.3% | #33 of 49, top 68% | max | Epoch AI | |
| GDP.pdf | 23% | #16 of 36, top 45% | max | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 31.4% | high | Epoch AI | ||
| ARC-AGI-2 | 4.6% | low | Epoch AI | ||
| ARC-AGI-2 | 59.3% | #31 of 83, top 38% | max | Epoch AI | |
| ARC-AGI-2 | 18.1% | medium | Epoch AI | ||
| ARC-AGI-2 | 0% | none | Epoch AI | ||
| ARC-AGI-2 | 41.9% | xhigh | Epoch AI | ||
| NYT Connections (extended) | 68.7% | #51 of 91, top 57% | high reasoning | Lech Mazur benchmarks | |
| ARC-AGI-1 | 70.3% | high | Epoch AI | ||
| ARC-AGI-1 | 37.7% | low | Epoch AI | ||
| ARC-AGI-1 | 86.7% | #33 of 83, top 40% | max | Epoch AI | |
| ARC-AGI-1 | 61% | medium | Epoch AI | ||
| ARC-AGI-1 | 8.8% | none | Epoch AI | ||
| ARC-AGI-1 | 73% | xhigh | Epoch AI | ||
| CritPt | 15.4% | high | Epoch AI | ||
| CritPt | 2.6% | low | Epoch AI | ||
| CritPt | 19.4% | #26 of 134, top 20% | max | Epoch AI | |
| CritPt | 10.6% | medium | Epoch AI | ||
| CritPt | 1.1% | none | Epoch AI | ||
| CritPt | 17.4% | xhigh | Epoch AI | ||
| Chess Puzzles | 31% | #34 of 129, top 27% | max | Epoch AI | 2026-09-22 |
| LMArena Hard Prompts | 1411 | #116 of 297, top 40% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 7% | #67 of 74, top 91% | max | Epoch AI | 2026-09-22 |
| DTBench | 90.1% | #38 of 151, top 26% | max | Epoch AI | |
| LMCA | 44.5% | #35 of 125, top 29% | max | Epoch AI | |
| Epoch Capabilities Index | 156.28 | #24 of 213, top 12% | Epoch AI | 2026-09-22 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 78.9% | #16 of 81, top 20% | max | Epoch AI | 2026-09-22 |
| FrontierMath Tier 4 | 56.1% | #17 of 63, top 27% | max | Epoch AI | 2026-09-22 |
| OTIS Mock AIME 2024-2025 | 98.9% | #18 of 173, top 11% | max | Epoch AI | 2026-09-22 |
| ProofBench | 64% | #17 of 77, top 23% | Epoch AI | ||
| LMArena Math | 1416 | #108 of 285, top 38% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 90.5% | #36 of 186, top 20% | max | Epoch AI | 2026-09-22 |
| SimpleQA Verified | 41.4% | #39 of 77, top 51% | max | Epoch AI | 2026-09-22 |
| LMArena Expert | 1444 | #76 of 273, top 28% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1217 | #72 of 122, top 60% | LMArena | 2026-10-09 | |
| Blueprint-Bench 2 | 31.2% | #14 of 31, top 46% | Epoch AI | ||
| Furniture Assembly | 44.2% | #12 of 31, top 39% | max | Epoch AI | 2026-09-28 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1386 | #117 of 297, top 40% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1433 | #115 of 285, top 41% | LMArena | 2026-10-08 | |
| LMArena French | 1420 | #100 of 223, top 45% | LMArena | 2026-10-08 | |
| LMArena German | 1369 | #117 of 231, top 51% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1369 | #87 of 211, top 42% | LMArena | 2026-10-08 | |
| LMArena Korean | 1360 | #93 of 213, top 44% | LMArena | 2026-10-08 | |
| LMArena Russian | 1394 | #113 of 283, top 40% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1393 | #117 of 226, top 52% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1409 | #88 of 298, top 30% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1409 | #108 of 291, top 38% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1391 | #128 of 297, top 44% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1363 | #118 of 295, top 40% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1396 | #124 of 295, top 43% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.10 | $0.50 | $0.01 | 2026-10-10 |
| bedrock | $0.10 | $0.50 | $0.01 | 2026-10-10 |
| openai | $0.10 | $0.50 | $0.01 | 2026-10-10 |
| openrouter | $0.10 | $0.50 | $0.01 | 2026-10-10 |
Compare GPT-6 Luna
- GPT-6 Luna vs GPT-5.6 Luna
- GPT-6 Luna vs DeepSeek V4 Flash
- GPT-6 Luna vs Grok 4.7
- GPT-6 Luna vs GPT-5.2
- GPT-6 Luna vs DeepSeek V4.1 Flash
- GPT-6 Luna vs Gemini 3.6 Flash
- GPT-6 Luna vs GPT-5.2 Pro
- GPT-6 Luna vs Claude Fable 5.1
- GPT-6 Luna vs Gemini 3.8 Flash
- GPT-6 Luna vs Kimi K3
- GPT-6 Luna vs Grok 4.6
- GPT-6 Luna vs Qwen3.8 Max
- GPT-6 Luna vs GLM-5.3
- GPT-6 Luna vs Muse Spark 1.3
Other OpenAI models
- GPT-6 Astra70.8
- GPT-6.1 Sol65.6
- GPT-5.6 Sol65.0
- GPT-5.5 Pro64.3
- GPT-5.563.4
- GPT-6 Sol61.8
- GPT-5.459.4
- GPT-5.6 Terra59.2
Frequently asked questions
How good is GPT-6 Luna?
GPT-6 Luna by OpenAI ranks 36th of 354 ranked models on the Noometry Index as of October 2026, with a score of 53.3. Its strongest category is math, where it ranks 15th. API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens, with a 1.05M-token context window.
How much does GPT-6 Luna cost?
GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens on OpenAI's own API, with cached input at $0.01.
What is GPT-6 Luna's context window?
GPT-6 Luna accepts up to 1.05M tokens of input and can write up to 128K tokens in one response.
Is GPT-6 Luna open source?
No. GPT-6 Luna is proprietary and available only through OpenAI's API and partner platforms.
What are GPT-6 Luna's strengths and weaknesses?
Relative to other ranked models, GPT-6 Luna places best in math, coding, reasoning and lowest in multilingual, writing & preference, long context.
What is GPT-6 Luna best at?
Its best category is math, where it ranks 15th on Noometry.