Model comparison

Gemini 2.0 Pro vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 39.1 on the Noometry Index.

Last verified . 2 shared benchmarks.

Gemini 2.0 Pro Google

39.1

Rank #173 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Gemini 2.0 Pro scores higher in 1 category and Grok 4.6 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 22.3.
  • The biggest single-benchmark swing is GPQA Diamond: 65.7% for Gemini 2.0 Pro and 94% for Grok 4.6.

Side by side

Gemini 2.0 Pro and Grok 4.6 specifications
Gemini 2.0 ProGrok 4.6
ProviderGooglexAI
Noometry Index39.156.9
Released2025-02-052026-08-12
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1449

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Gemini 2.0 Pro: 37.8 (#187), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
DeepSWE—67.5%
FrontierCode—48%
Aider Polyglot35.6%—
CursorBench—41.4%
LMArena WebDev—1617
FrontierSWE—25.3%
SciCode—56.5%
WeirdML—67.3%
LiveBench Coding63.5%—
LMArena Coding—1465
ALE-Bench—1,508

Agentic & Tool Use Not comparable

Gemini 2.0 Pro: —, Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
APEX-Agents—65.3%
GDP.pdf—17.2%
Vending-Bench 2—9,047

Reasoning Grok 4.6 leads

Gemini 2.0 Pro: 22.3 (#198), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
Epoch Capabilities Index135.06156.44
ARC-AGI-2—67.1%
SimpleBench—75.9%
NYT Connections (extended)—80%
ARC-AGI-1—87.5%
CritPt—19.7%
Chess Puzzles—40%
EnigmaEval0.7%—
EBR-Bench—30.5%
LiveBench Reasoning60.1%—
LMArena Hard Prompts—1447
Mystery Game Puzzles—34%
DTBench—97.3%
LiveBench Data Analysis68%—
LMCA—48.5%
LiveBench65.1%—

Math Grok 4.6 leads

Gemini 2.0 Pro: 39.7 (#100), Grok 4.6: 67.0 (#24)

Math benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
FrontierMath (Tiers 1-3)—66%
FrontierMath Tier 4—31.7%
OTIS Mock AIME 2024-2025—99.2%
ProofBench—51%
LiveBench Math71%—
LMArena Math—1423
MATH Level 583.5%—

Knowledge Grok 4.6 leads

Gemini 2.0 Pro: 36.5 (#167), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
GPQA Diamond65.7%94%
SimpleQA Verified—49.3%
Confabulations18.4%—
LMArena Expert—1467

Multimodal Not comparable

Gemini 2.0 Pro: —, Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
LMArena Vision—1263
Blueprint-Bench 2—33.2%
Furniture Assembly—40%
LMArena Document—1452

Multilingual Not comparable

Gemini 2.0 Pro: —, Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
LMArena Non-English—1420
LMArena Chinese—1480
LMArena French—1461
LMArena German—1431
LMArena Japanese—1376
LMArena Korean—1397
LMArena Russian—1422
LMArena Spanish—1404

Instruction Following Too close to call

Gemini 2.0 Pro: 75.5 (#59), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
LiveBench Instruction Following83.4%—
LMArena Instruction Following—1431

Long Context Grok 4.6 leads

Gemini 2.0 Pro: 29.2 (#292), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
Fiction.LiveBench41.7%—
LMArena Longer Query—1454

Writing & Preference Grok 4.6 leads

Gemini 2.0 Pro: 52.7 (#165), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkGemini 2.0 ProGrok 4.6
LMArena Text—1428
LMArena Creative Writing—1428
LMArena Multi-Turn—1425
LiveBench Language44.9%—

Frequently asked questions

Is Gemini 2.0 Pro better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 39.1 on the Noometry Index.

Is Gemini 2.0 Pro or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 37.8 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Pro and Grok 4.6 share?

2 benchmarks have published results for both models. Gemini 2.0 Pro has 14 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper