OpenAI, open weights

gpt-oss-20b

gpt-oss-20b by OpenAI ranks 255th of 354 ranked models on the Noometry Index as of October 2026, with a score of 32.5. Its strongest category is math, where it ranks 103rd. API pricing starts at $0.018 per million input tokens and $0.09 per million output tokens, with a 131K-token context window.

Last verified

Specifications

Noometry rank
#255 of 354
Index score
32.5
Evidence
Confirmed 34 results
Provider
OpenAI
Released
August 5, 2025
Weights
Open weights
Reasoning
Yes
Context window
131K
Max output
16K
Input price
$0.018 / M
Output price
$0.09 / M
Blended price
$0.036 / M
Output speed
96 tokens/s Kagi
Value
#1 of 219
Knowledge cutoff
Unknown
Input
text
Hugging Face
openai/gpt-oss-20b

Category scores

Each category score combines every public result we have in that category.

gpt-oss-20b category scores
  1. Coding 37.6
  2. Agentic & Tool Use 9.3
  3. Reasoning 19.3
  4. Math 39.4
  5. Knowledge 34.6
  6. Multilingual 42.2
  7. Instruction Following 61.8
  8. Long Context 37.9
  9. Writing & Preference 35.5
gpt-oss-20b category ranks
CategoryScoreRankResults
Coding37.6#1923
Agentic & Tool Use9.3#1541
Reasoning19.3#2616
Math39.4#1033
Knowledge34.6#1954
Multilingual42.2#1971
Instruction Following61.8#2402
Long Context37.9#2091
Writing & Preference35.5#2655

Strengths and weaknesses

Categories where gpt-oss-20b places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

gpt-oss-20b: strongest categories
CategoryScorevs medianRank
Math39.4+2.9#103 of 327, top 32%
Coding37.6−1.1#192 of 340, top 57%
Knowledge34.6−2.7#195 of 314, top 63%

Weakest categories

gpt-oss-20b: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use9.3−21.0#154 of 154, top 100%
Writing & Preference35.5−18.2#265 of 312, top 85%
Instruction Following61.8−9.5#240 of 305, top 79%

Closest competitors

The models ranked just above and below gpt-oss-20b. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to gpt-oss-20b
ModelRankScoreBlended $/MSpeed
Tulu 3 (Tülu 3) 70B#25133.0——Compare
DeepSeek-R1-Distill-Qwen-14B#25232.7——Compare
Qwen1.5-14B#25332.7——Compare
Olmo 2 0325 32b Instruct#25432.7——Compare
Laguna M.1#25632.5——Compare
Command R+#25732.4$4.38—Compare
Granite 3.1 8b Instruct#25832.4——Compare
Pixtral Large#25932.2$3—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

gpt-oss-20b Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SciCode34.4%#101 of 121, top 84%highEpoch AI
WeirdML40.9%#75 of 119, top 64%highEpoch AI
WeirdML36.8%mediumEpoch AI
LMArena Coding1306#196 of 294, top 67%LMArena2026-10-08
ALE-Bench566.05#83 of 105, top 80%Epoch AI

Agentic & Tool Use

gpt-oss-20b Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench3.4%#41 of 41, top 100%Epoch AI

Reasoning

gpt-oss-20b Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Kagi LLM Benchmark53.2%#56 of 99, top 57%Kagi LLM Benchmark
CritPt1.4%#76 of 134, top 57%highEpoch AI
Chess Puzzles0%highEpoch AI2026-08-27
Chess Puzzles2%lowEpoch AI2026-08-27
Chess Puzzles4%#98 of 129, top 76%mediumEpoch AI2026-08-27
LMArena Hard Prompts1274#202 of 297, top 69%LMArena2026-10-08
DTBench68%#97 of 151, top 65%highEpoch AI
LMCA14.5%#106 of 125, top 85%highEpoch AI
Epoch Capabilities Index137.82#116 of 213, top 55%Epoch AI2025-08-05

Math

gpt-oss-20b Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202550.8%highEpoch AI2026-08-27
OTIS Mock AIME 2024-202540.3%lowEpoch AI2026-08-27
OTIS Mock AIME 2024-202565.3%#100 of 173, top 58%mediumEpoch AI2026-08-27
Omni-MATH56.5%#11 of 57, top 20%HELM Capabilities
LMArena Math1317#172 of 285, top 61%LMArena2026-10-08

Knowledge

gpt-oss-20b Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond46%highEpoch AI2026-08-06
GPQA Diamond53.2%lowEpoch AI2026-08-27
GPQA Diamond60.8%#113 of 186, top 61%mediumEpoch AI2026-08-27
MMLU-Pro74%#28 of 58, top 49%HELM Capabilities
GPQA (HELM)59.4%#25 of 57, top 44%HELM Capabilities
LMArena Expert1258#191 of 273, top 70%LMArena2026-10-08

Multilingual

gpt-oss-20b Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1268#197 of 297, top 67%LMArena2026-10-08
LMArena Chinese1314#183 of 285, top 65%LMArena2026-10-08
LMArena German1255#173 of 231, top 75%LMArena2026-10-08
LMArena Japanese1244#142 of 211, top 68%LMArena2026-10-08
LMArena Korean1236#149 of 213, top 70%LMArena2026-10-08
LMArena Russian1278#192 of 283, top 68%LMArena2026-10-08
LMArena Spanish1267#175 of 226, top 78%LMArena2026-10-08

Instruction Following

gpt-oss-20b Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval73.2%#53 of 57, top 93%HELM Capabilities
LMArena Instruction Following1236#224 of 298, top 76%LMArena2026-10-08

Long Context

gpt-oss-20b Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1250#222 of 291, top 77%LMArena2026-10-08

Writing & Preference

gpt-oss-20b Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1287#200 of 297, top 68%LMArena2026-10-08
LMArena Creative Writing1201#232 of 295, top 79%LMArena2026-10-08
EQ-Bench Creative Writing666#112 of 115, top 98%EQ-Bench
WildBench73.7%#48 of 57, top 85%HELM Capabilities
LMArena Multi-Turn1268#212 of 295, top 72%LMArena2026-10-08

API pricing by provider

gpt-oss-20b API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$0.07$0.30—2026-10-10
deepinfra$0.03$0.14—2026-10-10
groq$0.075$0.30$0.03752026-10-10
openrouter$0.018$0.09$0.0092026-10-10
vertex$0.07$0.25$0.0072026-10-10

Compare gpt-oss-20b

Other OpenAI models

Frequently asked questions

How good is gpt-oss-20b?

gpt-oss-20b by OpenAI ranks 255th of 354 ranked models on the Noometry Index as of October 2026, with a score of 32.5. Its strongest category is math, where it ranks 103rd. API pricing starts at $0.018 per million input tokens and $0.09 per million output tokens, with a 131K-token context window.

How much does gpt-oss-20b cost?

gpt-oss-20b costs $0.018 per million input tokens and $0.09 per million output tokens on openrouter, with cached input at $0.009.

What is gpt-oss-20b's context window?

gpt-oss-20b accepts up to 131K tokens of input and can write up to 16K tokens in one response.

Is gpt-oss-20b open source?

Yes. gpt-oss-20b's weights are downloadable from Hugging Face (openai/gpt-oss-20b); check the license for commercial terms.

How fast is gpt-oss-20b?

gpt-oss-20b generated about 96 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are gpt-oss-20b's strengths and weaknesses?

Relative to other ranked models, gpt-oss-20b places best in math, coding, knowledge and lowest in agentic & tool use, writing & preference, instruction following.

What is gpt-oss-20b best at?

Its best category is math, where it ranks 103rd on Noometry.