Model comparison

GPT-4.1 vs gpt-oss-120b

GPT-4.1 and gpt-oss-120b score almost the same on the Noometry Index (35.9 vs 36.3), so choose on price, context window or the category you care about most.

Last verified . 37 shared benchmarks.

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Summary

  • They share 37 benchmarks with published results for both. GPT-4.1 scores higher in 6 categories and gpt-oss-120b in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-120b leads 52.5 to 22.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 38.3% for GPT-4.1 and 88.9% for gpt-oss-120b.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $2 / $8 for GPT-4.1.
  • GPT-4.1 accepts more context: 1.05M tokens versus 131K.
  • gpt-oss-120b has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 and gpt-oss-120b specifications
GPT-4.1gpt-oss-120b
ProviderOpenAIOpenAI
Noometry Index35.936.3
Released2025-04-142025-08-05
WeightsProprietaryOpen
Context window1.05M131K
Max output33K41K
Input $ / M tokens$2$0.037
Output $ / M tokens$8$0.17
Results tracked5248

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4.1: 34.4 (#238), gpt-oss-120b: 33.5 (#256)

Coding benchmarks
BenchmarkGPT-4.1gpt-oss-120b
SWE-bench Verified (bash only)39.6%26%
Aider Polyglot52.4%41.8%
WeirdML39%48.2%
LMArena Coding13911380
ALE-Bench558.1575.62
SWE-bench Verified48.5%—
SciCode—36%
CadEval42%—
AlgoTune—1.41

Agentic & Tool Use GPT-4.1 leads

GPT-4.1: 34.7 (#43), gpt-oss-120b: 12.2 (#153)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1gpt-oss-120b
Terminal-Bench—18.7%
APEX-Agents—4.4%
Berkeley Function Calling Leaderboard54%—
METR Time Horizons—56.6%
Vending-Bench 2—-21.53

Reasoning gpt-oss-120b leads

GPT-4.1: 11.7 (#339), gpt-oss-120b: 20.0 (#245)

Reasoning benchmarks
BenchmarkGPT-4.1gpt-oss-120b
SimpleBench27%22.1%
Kagi LLM Benchmark52.3%58.6%
Chess Puzzles6%20%
LMArena Hard Prompts13841364
DTBench68.3%76.3%
LMCA25.6%22.1%
Epoch Capabilities Index136.78139.93
ARC-AGI-20.4%—
ARC-AGI-15.5%—
CritPt—1.1%
EnigmaEval2.2%—
Mystery Game Puzzles—2%
Surface Evolver Bench—25%
ForecastBench61.5—

Math gpt-oss-120b leads

GPT-4.1: 22.3 (#280), gpt-oss-120b: 52.5 (#50)

Math benchmarks
BenchmarkGPT-4.1gpt-oss-120b
OTIS Mock AIME 2024-202538.3%88.9%
Omni-MATH47.1%68.8%
LMArena Math13701389
FrontierMath (Tiers 1-3)6%—
MATH Level 583%—
FrontierMath (Feb 2025 set)5.5%—
FrontierMath Tier 4 (v1)0%—

Knowledge gpt-oss-120b leads

GPT-4.1: 37.1 (#160), gpt-oss-120b: 42.4 (#96)

Knowledge benchmarks
BenchmarkGPT-4.1gpt-oss-120b
GPQA Diamond66.9%75.8%
MMLU-Pro81.1%79.5%
Vectara Hallucination Rate5.6%14.2%
GPQA (HELM)65.9%68.4%
LMArena Expert13641356
Humanity's Last Exam5.4%—
SimpleQA Verified31.1%—
Confabulations—15.7%

Multimodal Not comparable

GPT-4.1: 38.2 (#67), gpt-oss-120b: —

Multimodal benchmarks
BenchmarkGPT-4.1gpt-oss-120b
LMArena Vision1211—
GeoBench72%—

Multilingual GPT-4.1 leads

GPT-4.1: 49.4 (#133), gpt-oss-120b: 48.0 (#147)

Multilingual benchmarks
BenchmarkGPT-4.1gpt-oss-120b
LMArena Non-English13701351
LMArena Chinese13821385
LMArena French13821369
LMArena German13811353
LMArena Japanese13191331
LMArena Korean13391282
LMArena Russian13771343
LMArena Spanish13761389

Instruction Following GPT-4.1 leads

GPT-4.1: 71.3 (#153), gpt-oss-120b: 69.3 (#173)

Instruction Following benchmarks
BenchmarkGPT-4.1gpt-oss-120b
IFEval83.8%83.6%
LMArena Instruction Following13671318

Long Context GPT-4.1 leads

GPT-4.1: 40.0 (#163), gpt-oss-120b: 31.4 (#278)

Long Context benchmarks
BenchmarkGPT-4.1gpt-oss-120b
Fiction.LiveBench63.9%44.4%
LMArena Longer Query13851319

Writing & Preference GPT-4.1 leads

GPT-4.1: 57.6 (#125), gpt-oss-120b: 46.5 (#217)

Writing & Preference benchmarks
BenchmarkGPT-4.1gpt-oss-120b
LMArena Text13831365
LMArena Creative Writing13631275
EQ-Bench Creative Writing1420961
WildBench85.4%84.5%
LMArena Multi-Turn13981340
Short-Story Creative Writing—77.1%

Frequently asked questions

Is GPT-4.1 better than gpt-oss-120b?

GPT-4.1 and gpt-oss-120b score almost the same on the Noometry Index (35.9 vs 36.3), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-4.1 or gpt-oss-120b?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; GPT-4.1 lists at $2 and $8.

Is GPT-4.1 or gpt-oss-120b better for coding?

They score almost the same on coding (34.4 vs 33.5); test both on your own repository before choosing.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 131K.

How many benchmarks do GPT-4.1 and gpt-oss-120b share?

37 benchmarks have published results for both models. GPT-4.1 has 52 scored results on Noometry and gpt-oss-120b has 48.

Related comparisons

Go deeper