Model comparison

Gemini 1.5 Flash (May 2024) vs GPT-4o mini

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 25.5 on the Noometry Index.

Last verified . 40 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Summary

  • They share 40 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 8 categories and GPT-4o mini in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 1.5 Flash (May 2024) leads 21.7 to 8.7.
  • The biggest single-benchmark swing is WeirdML: 24.9% for Gemini 1.5 Flash (May 2024) and 11.8% for GPT-4o mini.

Side by side

Gemini 1.5 Flash (May 2024) and GPT-4o mini specifications
Gemini 1.5 Flash (May 2024)GPT-4o mini
ProviderGoogleOpenAI
Noometry Index33.225.5
Released2024-05-142024-07-18
WeightsProprietaryProprietary
Context window—128K
Max output—16K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked4260

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), GPT-4o mini: 22.0 (#335)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
WeirdML24.9%11.8%
BigCodeBench Instruct43.5%46.1%
LMArena Coding12611290
BigCodeBench Complete55.1%57.4%
HumanEval+75.6%83.5%
MBPP+67.5%72.2%
Aider Polyglot—3.6%
LiveBench Coding—43.1%

Agentic & Tool Use Too close to call

Gemini 1.5 Flash (May 2024): 26.6 (#102), GPT-4o mini: 27.5 (#101)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
BALROG14.6%17.4%

Reasoning Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), GPT-4o mini: 8.7 (#347)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
LMArena Hard Prompts12571267
DTBench53.8%54.4%
Epoch Capabilities Index129.36126.56
PIQA87.5%88.7%
ARC-AGI-2—0%
SimpleBench—10.7%
Kagi LLM Benchmark—28.8%
Chess Puzzles—0%
LiveBench Reasoning—32.8%
Mystery Game Puzzles—12%
LiveBench Data Analysis—50%
LMCA—10.4%
ForecastBench53.9—
LiveBench—41.3%

Math Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), GPT-4o mini: 10.4 (#314)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
OTIS Mock AIME 2024-202516.3%6.9%
Omni-MATH30.4%28%
LMArena Math12691267
MATH Level 561.9%52.6%
GSM8K82.4%91.3%
FrontierMath (Tiers 1-3)—0.7%
LiveBench Math—36.3%
FrontierMath (Feb 2025 set)0%—

Knowledge Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), GPT-4o mini: 17.7 (#284)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
GPQA Diamond47.3%37.7%
MMLU-Pro67.8%60.3%
GPQA (HELM)43.7%36.8%
LMArena Expert12331235
BoolQ85.8%88.7%
MMLU77.9%81.8%
SimpleQA Verified—8.3%
Confabulations—37.2%

Multimodal Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 36.0 (#81), GPT-4o mini: 25.9 (#122)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
LMArena Vision11411066
Video-MME70.3%64.8%
GeoBench76%64%
VPCT—34%

Multilingual Too close to call

Gemini 1.5 Flash (May 2024): 42.9 (#189), GPT-4o mini: 42.0 (#199)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
LMArena Non-English12781266
LMArena Chinese12951265
LMArena French12581297
LMArena German12621272
LMArena Japanese12521216
LMArena Korean12211195
LMArena Russian12881275
LMArena Spanish12431276

Instruction Following Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 66.8 (#205), GPT-4o mini: 61.9 (#239)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
IFEval83.1%78.2%
LMArena Instruction Following12581258
LiveBench Instruction Following—56.8%

Long Context Too close to call

Gemini 1.5 Flash (May 2024): 39.0 (#187), GPT-4o mini: 39.1 (#186)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
LMArena Longer Query12841289

Writing & Preference Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 48.7 (#196), GPT-4o mini: 39.5 (#248)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)GPT-4o mini
LMArena Text12871286
LMArena Creative Writing12851268
WildBench79.2%79.1%
LMArena Multi-Turn12531285
Short-Story Creative Writing—67.2%
EQ-Bench Creative Writing—873
LiveBench Language—28.6%

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than GPT-4o mini?

Gemini 1.5 Flash (May 2024) is the stronger model overall, scoring 33.2 to 25.5 on the Noometry Index.

Is Gemini 1.5 Flash (May 2024) or GPT-4o mini better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 22.0 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and GPT-4o mini share?

40 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and GPT-4o mini has 60.

Related comparisons

Go deeper