Model comparison

Gemini 1.5 Flash (May 2024) vs Grok-2 (Dec 2024)

Gemini 1.5 Flash (May 2024) and Grok-2 (Dec 2024) score almost the same on the Noometry Index (33.2 vs 33.7), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

Gemini 1.5 Flash (May 2024) Google

33.2

Rank #246 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemini 1.5 Flash (May 2024) scores higher in 5 categories and Grok-2 (Dec 2024) in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 1.5 Flash (May 2024) leads 21.7 to 16.9.
  • The biggest single-benchmark swing is DTBench: 53.8% for Gemini 1.5 Flash (May 2024) and 65.2% for Grok-2 (Dec 2024).

Side by side

Gemini 1.5 Flash (May 2024) and Grok-2 (Dec 2024) specifications
Gemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
ProviderGooglexAI
Noometry Index33.233.7
Released2024-05-142024-08-13
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4234

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 34.4 (#236), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
WeirdML24.9%22.2%
LMArena Coding12611287
BigCodeBench Instruct43.5%—
LiveBench Coding—46.4%
BigCodeBench Complete55.1%—
HumanEval+75.6%—
MBPP+67.5%—

Agentic & Tool Use Not comparable

Gemini 1.5 Flash (May 2024): 26.6 (#102), Grok-2 (Dec 2024): —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
BALROG14.6%—

Reasoning Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 21.7 (#215), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
LMArena Hard Prompts12571272
DTBench53.8%65.2%
Epoch Capabilities Index129.36130.48
SimpleBench—22.7%
LiveBench Reasoning—54.8%
LiveBench Data Analysis—54.5%
ForecastBench53.9—
LiveBench—54.3%
PIQA87.5%—

Math Gemini 1.5 Flash (May 2024) leads

Gemini 1.5 Flash (May 2024): 22.1 (#281), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
OTIS Mock AIME 2024-202516.3%11.5%
LMArena Math12691283
MATH Level 561.9%63.5%
FrontierMath (Feb 2025 set)0%0.7%
Omni-MATH30.4%—
LiveBench Math—54.9%
GSM8K82.4%—

Knowledge Grok-2 (Dec 2024) leads

Gemini 1.5 Flash (May 2024): 26.2 (#260), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
GPQA Diamond47.3%53.8%
LMArena Expert12331254
MMLU-Pro67.8%—
Confabulations—20.1%
GPQA (HELM)43.7%—
BoolQ85.8%—
MMLU77.9%—

Multimodal Not comparable

Gemini 1.5 Flash (May 2024): 36.0 (#81), Grok-2 (Dec 2024): —

Multimodal benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
LMArena Vision1141—
Video-MME70.3%—
GeoBench76%—

Multilingual Too close to call

Gemini 1.5 Flash (May 2024): 42.9 (#189), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
LMArena Non-English12781282
LMArena Chinese12951289
LMArena French12581318
LMArena German12621287
LMArena Japanese12521244
LMArena Korean12211237
LMArena Russian12881286
LMArena Spanish12431281

Instruction Following Too close to call

Gemini 1.5 Flash (May 2024): 66.8 (#205), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
LMArena Instruction Following12581270
LiveBench Instruction Following—69.6%
IFEval83.1%—

Long Context Too close to call

Gemini 1.5 Flash (May 2024): 39.0 (#187), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
LMArena Longer Query12841276

Writing & Preference Too close to call

Gemini 1.5 Flash (May 2024): 48.7 (#196), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash (May 2024)Grok-2 (Dec 2024)
LMArena Text12871305
LMArena Creative Writing12851284
LMArena Multi-Turn12531290
Short-Story Creative Writing—63.6%
WildBench79.2%—
LiveBench Language—45.6%

Frequently asked questions

Is Gemini 1.5 Flash (May 2024) better than Grok-2 (Dec 2024)?

Gemini 1.5 Flash (May 2024) and Grok-2 (Dec 2024) score almost the same on the Noometry Index (33.2 vs 33.7), so choose on price, context window or the category you care about most.

Is Gemini 1.5 Flash (May 2024) or Grok-2 (Dec 2024) better for coding?

Gemini 1.5 Flash (May 2024) scores higher on coding benchmarks: 34.4 versus 33.3 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash (May 2024) and Grok-2 (Dec 2024) share?

24 benchmarks have published results for both models. Gemini 1.5 Flash (May 2024) has 42 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper