Model comparison

GPT-4o mini vs Mistral Nemo

GPT-4o mini and Mistral Nemo score almost the same on the Noometry Index (25.5 vs 26.4), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 9 benchmarks with published results for both. GPT-4o mini scores higher in 3 categories and Mistral Nemo in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Nemo leads 25.5 to 10.4.
  • The biggest single-benchmark swing is MATH Level 5: 52.6% for GPT-4o mini and 10.8% for Mistral Nemo.
  • Mistral Nemo is cheaper at $0.15 / $0.15 per million input/output tokens, against $0.15 / $0.60 for GPT-4o mini.
  • Mistral Nemo has downloadable open weights; the other is API-only.

Side by side

GPT-4o mini and Mistral Nemo specifications
GPT-4o miniMistral Nemo
ProviderOpenAIMistral AI
Noometry Index25.526.4
Released2024-07-182024-07-01
WeightsProprietaryOpen
Context window128K128K
Max output16K128K
Input $ / M tokens$0.15$0.15
Output $ / M tokens$0.60$0.15
Results tracked6010

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

GPT-4o mini: 22.0 (#335), Mistral Nemo: —

Coding benchmarks
BenchmarkGPT-4o miniMistral Nemo
Aider Polyglot3.6%—
WeirdML11.8%—
BigCodeBench Instruct46.1%—
LiveBench Coding43.1%—
LMArena Coding1290—
BigCodeBench Complete57.4%—
HumanEval+83.5%—
MBPP+72.2%—

Agentic & Tool Use GPT-4o mini leads

GPT-4o mini: 27.5 (#101), Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkGPT-4o miniMistral Nemo
BALROG17.4%17.6%
Berkeley Function Calling Leaderboard—27.6%

Reasoning Mistral Nemo leads

GPT-4o mini: 8.7 (#347), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkGPT-4o miniMistral Nemo
DTBench54.4%48.6%
Epoch Capabilities Index126.56118.68
PIQA88.7%83.5%
ARC-AGI-20%—
SimpleBench10.7%—
Kagi LLM Benchmark28.8%—
Chess Puzzles0%—
LiveBench Reasoning32.8%—
LMArena Hard Prompts1267—
Mystery Game Puzzles12%—
LiveBench Data Analysis50%—
LMCA10.4%—
LiveBench41.3%—

Math Mistral Nemo leads

GPT-4o mini: 10.4 (#314), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkGPT-4o miniMistral Nemo
MATH Level 552.6%10.8%
GSM8K91.3%84.2%
FrontierMath (Tiers 1-3)0.7%—
OTIS Mock AIME 2024-20256.9%—
Omni-MATH28%—
LiveBench Math36.3%—
LMArena Math1267—

Knowledge GPT-4o mini leads

GPT-4o mini: 17.7 (#284), Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkGPT-4o miniMistral Nemo
GPQA Diamond37.7%29.9%
BoolQ88.7%82.5%
SimpleQA Verified8.3%—
MMLU-Pro60.3%—
Confabulations37.2%—
GPQA (HELM)36.8%—
LMArena Expert1235—
MMLU81.8%—

Multimodal Not comparable

GPT-4o mini: 25.9 (#122), Mistral Nemo: —

Multimodal benchmarks
BenchmarkGPT-4o miniMistral Nemo
LMArena Vision1066—
Video-MME64.8%—
GeoBench64%—
VPCT34%—

Multilingual Not comparable

GPT-4o mini: 42.0 (#199), Mistral Nemo: —

Multilingual benchmarks
BenchmarkGPT-4o miniMistral Nemo
LMArena Non-English1266—
LMArena Chinese1265—
LMArena French1297—
LMArena German1272—
LMArena Japanese1216—
LMArena Korean1195—
LMArena Russian1275—
LMArena Spanish1276—

Instruction Following Not comparable

GPT-4o mini: 61.9 (#239), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkGPT-4o miniMistral Nemo
LiveBench Instruction Following56.8%—
IFEval78.2%—
LMArena Instruction Following1258—

Long Context Not comparable

GPT-4o mini: 39.1 (#186), Mistral Nemo: —

Long Context benchmarks
BenchmarkGPT-4o miniMistral Nemo
LMArena Longer Query1289—

Writing & Preference GPT-4o mini leads

GPT-4o mini: 39.5 (#248), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkGPT-4o miniMistral Nemo
EQ-Bench Creative Writing873881
LMArena Text1286—
LMArena Creative Writing1268—
Short-Story Creative Writing67.2%—
WildBench79.1%—
LMArena Multi-Turn1285—
LiveBench Language28.6%—

Frequently asked questions

Is GPT-4o mini better than Mistral Nemo?

GPT-4o mini and Mistral Nemo score almost the same on the Noometry Index (25.5 vs 26.4), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-4o mini or Mistral Nemo?

Mistral Nemo is cheaper. It lists at $0.15 per million input tokens and $0.15 per million output tokens; GPT-4o mini lists at $0.15 and $0.60.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do GPT-4o mini and Mistral Nemo share?

9 benchmarks have published results for both models. GPT-4o mini has 60 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper