Model comparison

Gemini 1.5 Flash 8B vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 29.9 on the Noometry Index.

Last verified . 21 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 0 categories and Grok 4.6 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.6 leads 67.0 to 14.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.6% for Gemini 1.5 Flash 8B and 99.2% for Grok 4.6.

Side by side

Gemini 1.5 Flash 8B and Grok 4.6 specifications
Gemini 1.5 Flash 8BGrok 4.6
ProviderGooglexAI
Noometry Index29.956.9
Released2024-10-032026-08-12
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked2149

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Gemini 1.5 Flash 8B: 35.5 (#225), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
LMArena Coding12181465
DeepSWE—67.5%
FrontierCode—48%
CursorBench—41.4%
LMArena WebDev—1617
FrontierSWE—25.3%
SciCode—56.5%
WeirdML—67.3%
ALE-Bench—1,508

Agentic & Tool Use Not comparable

Gemini 1.5 Flash 8B: —, Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
APEX-Agents—65.3%
GDP.pdf—17.2%
Vending-Bench 2—9,047

Reasoning Grok 4.6 leads

Gemini 1.5 Flash 8B: 20.0 (#244), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
LMArena Hard Prompts12091447
DTBench50%97.3%
ARC-AGI-2—67.1%
SimpleBench—75.9%
NYT Connections (extended)—80%
ARC-AGI-1—87.5%
CritPt—19.7%
Chess Puzzles—40%
EBR-Bench—30.5%
Mystery Game Puzzles—34%
LMCA—48.5%
Epoch Capabilities Index—156.44

Math Grok 4.6 leads

Gemini 1.5 Flash 8B: 14.2 (#302), Grok 4.6: 67.0 (#24)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
OTIS Mock AIME 2024-20254.6%99.2%
LMArena Math12071423
FrontierMath (Tiers 1-3)—66%
FrontierMath Tier 4—31.7%
ProofBench—51%

Knowledge Grok 4.6 leads

Gemini 1.5 Flash 8B: 16.0 (#289), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
GPQA Diamond33%94%
LMArena Expert11851467
SimpleQA Verified—49.3%

Multimodal Grok 4.6 leads

Gemini 1.5 Flash 8B: 28.2 (#115), Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
LMArena Vision10441263
Blueprint-Bench 2—33.2%
Furniture Assembly—40%
LMArena Document—1452

Multilingual Grok 4.6 leads

Gemini 1.5 Flash 8B: 38.5 (#229), Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
LMArena Non-English12151420
LMArena Chinese12311480
LMArena French12341461
LMArena German12061431
LMArena Japanese11501376
LMArena Korean11401397
LMArena Russian12361422
LMArena Spanish12121404

Instruction Following Grok 4.6 leads

Gemini 1.5 Flash 8B: 62.8 (#236), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
LMArena Instruction Following11991431

Long Context Grok 4.6 leads

Gemini 1.5 Flash 8B: 37.0 (#225), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
LMArena Longer Query12191454

Writing & Preference Grok 4.6 leads

Gemini 1.5 Flash 8B: 42.8 (#232), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.6
LMArena Text12261428
LMArena Creative Writing12181428
LMArena Multi-Turn11851425

Frequently asked questions

Is Gemini 1.5 Flash 8B better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 29.9 on the Noometry Index.

Is Gemini 1.5 Flash 8B or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 35.5 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and Grok 4.6 share?

21 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper