Model comparison

Gemini 3.1 Flash Lite vs Grok-3 mini

Gemini 3.1 Flash Lite and Grok-3 mini score almost the same on the Noometry Index (40.8 vs 41.2), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

Gemini 3.1 Flash Lite Google

40.8

Rank #144 Confirmed

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Gemini 3.1 Flash Lite scores higher in 4 categories and Grok-3 mini in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.1 Flash Lite leads 22.9 to 13.6.
  • The biggest single-benchmark swing is WeirdML: 52.2% for Gemini 3.1 Flash Lite and 42.6% for Grok-3 mini.

Side by side

Gemini 3.1 Flash Lite and Grok-3 mini specifications
Gemini 3.1 Flash LiteGrok-3 mini
ProviderGooglexAI
Noometry Index40.841.2
Released2026-03-032025-04-09
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$0.25—
Output $ / M tokens$1.50—
Results tracked3835

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok-3 mini leads

Gemini 3.1 Flash Lite: 37.8 (#188), Grok-3 mini: 40.8 (#131)

Coding benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
WeirdML52.2%42.6%
LMArena Coding14001379
Aider Polyglot—49.3%
LMArena WebDev1256—
SciCode41.9%—
ALE-Bench797.73—

Agentic & Tool Use Not comparable

Gemini 3.1 Flash Lite: 30.2 (#79), Grok-3 mini: —

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
DeepResearch Bench37.3%—

Reasoning Gemini 3.1 Flash Lite leads

Gemini 3.1 Flash Lite: 22.9 (#186), Grok-3 mini: 13.6 (#334)

Reasoning benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
Kagi LLM Benchmark67.2%61.3%
LMArena Hard Prompts14071375
Epoch Capabilities Index144.47140.35
ARC-AGI-2—0.4%
NYT Connections (extended)8.2%—
ARC-AGI-1—16.5%
CritPt1.1%—
Chess Puzzles25%—
EnigmaEval3%—
Thematic Generalization63.3%—
DTBench76.8%—
LMCA35%—
ForecastBench54.4—

Math Grok-3 mini leads

Gemini 3.1 Flash Lite: 40.7 (#90), Grok-3 mini: 42.1 (#85)

Math benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
OTIS Mock AIME 2024-202580%77.8%
LMArena Math14281386
FrontierMath (Tiers 1-3)27.7%—
Omni-MATH—31.8%
MATH Level 5—90.9%
FrontierMath (Feb 2025 set)—5.9%

Knowledge Grok-3 mini leads

Gemini 3.1 Flash Lite: 41.9 (#104), Grok-3 mini: 46.4 (#81)

Knowledge benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
GPQA Diamond81.8%76.3%
LMArena Expert13981395
Humanity's Last Exam8.6%—
MMLU-Pro—79.9%
Confabulations—10.8%
Vectara Hallucination Rate8.2%—
GPQA (HELM)—67.5%

Multimodal Not comparable

Gemini 3.1 Flash Lite: 39.4 (#60), Grok-3 mini: —

Multimodal benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
LMArena Vision1240—

Multilingual Gemini 3.1 Flash Lite leads

Gemini 3.1 Flash Lite: 52.3 (#86), Grok-3 mini: 48.1 (#145)

Multilingual benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
LMArena Non-English14111352
LMArena Chinese14611387
LMArena French14241357
LMArena German14291349
LMArena Japanese14131342
LMArena Korean13921335
LMArena Russian14201353
LMArena Spanish14211381

Instruction Following Grok-3 mini leads

Gemini 3.1 Flash Lite: 72.7 (#131), Grok-3 mini: 78.5 (#9)

Instruction Following benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
LMArena Instruction Following13771357
IFEval—95.1%

Long Context Gemini 3.1 Flash Lite leads

Gemini 3.1 Flash Lite: 42.5 (#122), Grok-3 mini: 41.0 (#147)

Long Context benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
LMArena Longer Query13941372
Fiction.LiveBench—66.7%

Writing & Preference Gemini 3.1 Flash Lite leads

Gemini 3.1 Flash Lite: 60.9 (#94), Grok-3 mini: 52.5 (#169)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Flash LiteGrok-3 mini
LMArena Text14161370
LMArena Creative Writing14011342
LMArena Multi-Turn14171355
Short-Story Creative Writing—73.5%
WildBench—65.1%

Frequently asked questions

Is Gemini 3.1 Flash Lite better than Grok-3 mini?

Gemini 3.1 Flash Lite and Grok-3 mini score almost the same on the Noometry Index (40.8 vs 41.2), so choose on price, context window or the category you care about most.

Is Gemini 3.1 Flash Lite or Grok-3 mini better for coding?

Grok-3 mini scores higher on coding benchmarks: 40.8 versus 37.8 in the Noometry coding category.

How many benchmarks do Gemini 3.1 Flash Lite and Grok-3 mini share?

22 benchmarks have published results for both models. Gemini 3.1 Flash Lite has 38 scored results on Noometry and Grok-3 mini has 35.

Related comparisons

Go deeper