OpenAI, open weights

gpt-oss-120b

gpt-oss-120b by OpenAI ranks 217th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.3. Its strongest category is math, where it ranks 50th. API pricing starts at $0.037 per million input tokens and $0.17 per million output tokens, with a 131K-token context window.

Last verified

Specifications

Noometry rank
#217 of 354
Index score
36.3
Evidence
Confirmed 48 results
Provider
OpenAI
Released
August 5, 2025
Weights
Open weights
Reasoning
Yes
Context window
131K
Max output
41K
Input price
$0.037 / M
Output price
$0.17 / M
Blended price
$0.0703 / M
Output speed
55 tokens/s Kagi
Value
#6 of 219
Knowledge cutoff
August 2025
Input
text

Category scores

Each category score combines every public result we have in that category.

gpt-oss-120b category scores
  1. Coding 33.5
  2. Agentic & Tool Use 12.2
  3. Reasoning 20.0
  4. Math 52.5
  5. Knowledge 42.4
  6. Multilingual 48.0
  7. Instruction Following 69.3
  8. Long Context 31.4
  9. Writing & Preference 46.5
gpt-oss-120b category ranks
CategoryScoreRankResults
Coding33.5#2565
Agentic & Tool Use12.2#1532
Reasoning20.0#2459
Math52.5#503
Knowledge42.4#966
Multilingual48.0#1471
Instruction Following69.3#1732
Long Context31.4#2782
Writing & Preference46.5#2176

Strengths and weaknesses

Categories where gpt-oss-120b places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

gpt-oss-120b: strongest categories
CategoryScorevs medianRank
Math52.5+15.9#50 of 327, top 16%
Knowledge42.4+5.0#96 of 314, top 31%
Multilingual48.0+0.6#147 of 297, top 50%

Weakest categories

gpt-oss-120b: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use12.2−18.2#153 of 154, top 100%
Long Context31.4−9.5#278 of 296, top 94%
Coding33.5−5.3#256 of 340, top 76%

Closest competitors

The models ranked just above and below gpt-oss-120b. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to gpt-oss-120b
ModelRankScoreBlended $/MSpeed
Llama 3.1 Nemotron Ultra 253b v1#21336.7——Compare
Granite 4.0 H Small#21436.5——Compare
Command A#21536.5$4.3828Compare
Grok Build 0.1#21636.4$1.25—Compare
Mistral Medium#21836.3$368Compare
GPT-4.1#21935.9$3.50116Compare
Deepseek Coder v2#22035.9——Compare
C4ai Aya Expanse 32b#22135.9——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

gpt-oss-120b Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)26%#33 of 39, top 85%SWE-bench2025-08-07
Aider Polyglot41.8%#26 of 44, top 60%highEpoch AI
Aider Polyglot41.8%#26 of 44, top 60%highEpoch AI
SciCode34%highEpoch AI
SciCode36%#93 of 121, top 77%lowEpoch AI
WeirdML48.2%#51 of 119, top 43%highEpoch AI
WeirdML48.2%#51 of 119, top 43%highEpoch AI
WeirdML41.9%mediumEpoch AI
LMArena Coding1380#150 of 294, top 52%LMArena2026-10-08
ALE-Bench575.62#82 of 105, top 79%Epoch AI
AlgoTune1.41#14 of 18, top 78%highEpoch AI

Agentic & Tool Use

gpt-oss-120b Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench18.7%#38 of 41, top 93%Epoch AI
APEX-Agents4.4%#49 of 49, top 100%Epoch AI
METR Time Horizons56.6%#20 of 32, top 63%Epoch AI
Vending-Bench 2-21.53#58 of 60, top 97%Epoch AI

Reasoning

gpt-oss-120b Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench22.1%#72 of 77, top 94%Epoch AI
Kagi LLM Benchmark58.6%#43 of 99, top 44%Kagi LLM Benchmark
CritPt1.1%#81 of 134, top 61%highEpoch AI
CritPt0%lowEpoch AI
Chess Puzzles20%#58 of 129, top 45%highEpoch AI2025-12-11
LMArena Hard Prompts1364#151 of 297, top 51%LMArena2026-10-08
Mystery Game Puzzles0%highEpoch AI2026-08-27
Mystery Game Puzzles2%#74 of 74, top 100%mediumEpoch AI2026-08-27
DTBench76.3%#85 of 151, top 57%highEpoch AI
LMCA22.1%#91 of 125, top 73%highEpoch AI
Surface Evolver Bench25%#23 of 25, top 92%Epoch AI
Epoch Capabilities Index139.93#108 of 213, top 51%Epoch AI2025-08-05

Math

gpt-oss-120b Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202588.9%#55 of 173, top 32%highEpoch AI2025-12-11
Omni-MATH68.8%#5 of 57, top 9%HELM Capabilities
LMArena Math1389#138 of 285, top 49%LMArena2026-10-08

Knowledge

gpt-oss-120b Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond75.8%#92 of 186, top 50%highEpoch AI2025-12-11
MMLU-Pro79.5%#17 of 58, top 30%HELM Capabilities
Confabulations (lower is better)15.7%#21 of 51, top 42%medium reasoningLech Mazur benchmarks
Vectara Hallucination Rate (lower is better)14.2%#83 of 96, top 87%Vectara Hallucination Leaderboard
GPQA (HELM)68.4%#12 of 57, top 22%HELM Capabilities
LMArena Expert1356#152 of 273, top 56%LMArena2026-10-08

Multilingual

gpt-oss-120b Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1351#147 of 297, top 50%LMArena2026-10-08
LMArena Chinese1385#147 of 285, top 52%LMArena2026-10-08
LMArena French1369#136 of 223, top 61%LMArena2026-10-08
LMArena German1353#128 of 231, top 56%LMArena2026-10-08
LMArena Japanese1331#109 of 211, top 52%LMArena2026-10-08
LMArena Korean1282#138 of 213, top 65%LMArena2026-10-08
LMArena Russian1343#149 of 283, top 53%LMArena2026-10-08
LMArena Spanish1389#121 of 226, top 54%LMArena2026-10-08

Instruction Following

gpt-oss-120b Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval83.6%#27 of 57, top 48%HELM Capabilities
LMArena Instruction Following1318#164 of 298, top 56%LMArena2026-10-08

Long Context

gpt-oss-120b Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench36.1%Epoch AI
Fiction.LiveBench44.4%#41 of 47, top 88%highEpoch AI
LMArena Longer Query1319#173 of 291, top 60%LMArena2026-10-08

Writing & Preference

gpt-oss-120b Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1365#147 of 297, top 50%LMArena2026-10-08
LMArena Creative Writing1275#187 of 295, top 64%LMArena2026-10-08
Short-Story Creative Writing77.1%#18 of 39, top 47%Epoch AI
EQ-Bench Creative Writing961#97 of 115, top 85%EQ-Bench
WildBench84.5%#14 of 57, top 25%HELM Capabilities
LMArena Multi-Turn1340#162 of 295, top 55%LMArena2026-10-08

API pricing by provider

gpt-oss-120b API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$0.15$0.60—2026-10-10
cerebras$0.35$0.75—2026-10-10
deepinfra$0.037$0.17—2026-10-10
fireworks$0.15$0.60$0.0152026-10-10
groq$0.15$0.60$0.0752026-10-10
openrouter$0.037$0.17—2026-10-10
together$0.15$0.60—2026-10-10
vertex$0.09$0.36—2026-10-10

Compare gpt-oss-120b

Other OpenAI models

Frequently asked questions

How good is gpt-oss-120b?

gpt-oss-120b by OpenAI ranks 217th of 354 ranked models on the Noometry Index as of October 2026, with a score of 36.3. Its strongest category is math, where it ranks 50th. API pricing starts at $0.037 per million input tokens and $0.17 per million output tokens, with a 131K-token context window.

How much does gpt-oss-120b cost?

gpt-oss-120b costs $0.037 per million input tokens and $0.17 per million output tokens on deepinfra.

What is gpt-oss-120b's context window?

gpt-oss-120b accepts up to 131K tokens of input and can write up to 41K tokens in one response.

Is gpt-oss-120b open source?

Yes. gpt-oss-120b's weights are downloadable from Hugging Face (openai/gpt-oss-120b); check the license for commercial terms.

How fast is gpt-oss-120b?

gpt-oss-120b generated about 55 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are gpt-oss-120b's strengths and weaknesses?

Relative to other ranked models, gpt-oss-120b places best in math, knowledge, multilingual and lowest in agentic & tool use, long context, coding.

What is gpt-oss-120b best at?

Its best category is math, where it ranks 50th on Noometry.