Model comparison

gpt-oss-120b vs Mistral Medium

gpt-oss-120b and Mistral Medium score almost the same on the Noometry Index (36.3 vs 36.3), so choose on price, context window or the category you care about most.

Last verified . 29 shared benchmarks.

gpt-oss-120b OpenAI

36.3

Rank #217 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 29 benchmarks with published results for both. gpt-oss-120b scores higher in 2 categories and Mistral Medium in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where gpt-oss-120b leads 52.5 to 28.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for gpt-oss-120b and 32.2% for Mistral Medium.
  • gpt-oss-120b is cheaper at $0.037 / $0.17 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • Mistral Medium accepts more context: 262K tokens versus 131K.

Side by side

gpt-oss-120b and Mistral Medium specifications
gpt-oss-120bMistral Medium
ProviderOpenAIMistral AI
Noometry Index36.336.3
Released2025-08-052023-12-11
WeightsOpenOpen
Context window131K262K
Max output41K262K
Input $ / M tokens$0.037$1.50
Output $ / M tokens$0.17$7.50
Results tracked4836

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

gpt-oss-120b: 33.5 (#256), Mistral Medium: 34.2 (#243)

Coding benchmarks
Benchmarkgpt-oss-120bMistral Medium
SciCode36%40.2%
WeirdML48.2%43.7%
LMArena Coding13801434
ALE-Bench575.62763.98
FrontierCode—8%
SWE-bench Verified (bash only)26%—
Aider Polyglot41.8%—
AlgoTune1.41—

Agentic & Tool Use Mistral Medium leads

gpt-oss-120b: 12.2 (#153), Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
Benchmarkgpt-oss-120bMistral Medium
Terminal-Bench18.7%—
APEX-Agents4.4%—
Berkeley Function Calling Leaderboard—37.7%
METR Time Horizons56.6%—
Vending-Bench 2-21.53—

Reasoning Mistral Medium leads

gpt-oss-120b: 20.0 (#245), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
Benchmarkgpt-oss-120bMistral Medium
Kagi LLM Benchmark58.6%50%
CritPt1.1%0%
LMArena Hard Prompts13641426
DTBench76.3%75.5%
LMCA22.1%26.1%
Surface Evolver Bench25%26.9%
SimpleBench22.1%—
Chess Puzzles20%—
Mystery Game Puzzles2%—
Epoch Capabilities Index139.93—

Math gpt-oss-120b leads

gpt-oss-120b: 52.5 (#50), Mistral Medium: 28.1 (#245)

Math benchmarks
Benchmarkgpt-oss-120bMistral Medium
OTIS Mock AIME 2024-202588.9%32.2%
LMArena Math13891408
ProofBench—9%
Omni-MATH68.8%—
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge gpt-oss-120b leads

gpt-oss-120b: 42.4 (#96), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
Benchmarkgpt-oss-120bMistral Medium
GPQA Diamond75.8%59.5%
Vectara Hallucination Rate14.2%22.7%
LMArena Expert13561408
Humanity's Last Exam—4.5%
MMLU-Pro79.5%—
Confabulations15.7%—
GPQA (HELM)68.4%—

Multimodal Not comparable

gpt-oss-120b: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
Benchmarkgpt-oss-120bMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

gpt-oss-120b: 48.0 (#147), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
Benchmarkgpt-oss-120bMistral Medium
LMArena Non-English13511408
LMArena Chinese13851447
LMArena French13691459
LMArena German13531432
LMArena Japanese13311378
LMArena Korean12821380
LMArena Russian13431411
LMArena Spanish13891433

Instruction Following Mistral Medium leads

gpt-oss-120b: 69.3 (#173), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
Benchmarkgpt-oss-120bMistral Medium
LMArena Instruction Following13181398
IFEval83.6%—

Long Context Mistral Medium leads

gpt-oss-120b: 31.4 (#278), Mistral Medium: 42.9 (#114)

Long Context benchmarks
Benchmarkgpt-oss-120bMistral Medium
LMArena Longer Query13191406
Fiction.LiveBench44.4%—

Writing & Preference Mistral Medium leads

gpt-oss-120b: 46.5 (#217), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
Benchmarkgpt-oss-120bMistral Medium
LMArena Text13651424
LMArena Creative Writing12751391
Short-Story Creative Writing77.1%77.3%
LMArena Multi-Turn13401418
EQ-Bench Creative Writing961—
WildBench84.5%—

Frequently asked questions

Is gpt-oss-120b better than Mistral Medium?

gpt-oss-120b and Mistral Medium score almost the same on the Noometry Index (36.3 vs 36.3), so choose on price, context window or the category you care about most.

Which is cheaper, gpt-oss-120b or Mistral Medium?

gpt-oss-120b is cheaper. It lists at $0.037 per million input tokens and $0.17 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is gpt-oss-120b or Mistral Medium better for coding?

They score almost the same on coding (33.5 vs 34.2); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 131K.

How many benchmarks do gpt-oss-120b and Mistral Medium share?

29 benchmarks have published results for both models. gpt-oss-120b has 48 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper