Model comparison

Gemini 3 Flash Preview vs Grok 4.7

Gemini 3 Flash Preview and Grok 4.7 score almost the same on the Noometry Index (52.3 vs 53.1), so choose on price, context window or the category you care about most.

Last verified . 31 shared benchmarks.

Gemini 3 Flash Preview Google

52.3

Rank #40 Confirmed

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Gemini 3 Flash Preview scores higher in 6 categories and Grok 4.7 in 4 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in multimodal, where Gemini 3 Flash Preview leads 45.5 to 35.5.
  • The biggest single-benchmark swing is Blueprint-Bench 2: 0% for Gemini 3 Flash Preview and 32.5% for Grok 4.7.
  • Gemini 3 Flash Preview is cheaper at $0.50 / $3 per million input/output tokens, against $2 / $6 for Grok 4.7.
  • Gemini 3 Flash Preview accepts more context: 1.05M tokens versus 500K.

Side by side

Gemini 3 Flash Preview and Grok 4.7 specifications
Gemini 3 Flash PreviewGrok 4.7
ProviderGooglexAI
Noometry Index52.353.1
Released2025-12-172026-09-21
WeightsProprietaryProprietary
Context window1.05M500K
Max output66K500K
Input $ / M tokens$0.50$2
Output $ / M tokens$3$6
Results tracked5939

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.7 leads

Gemini 3 Flash Preview: 50.9 (#42), Grok 4.7: 58.0 (#18)

Coding benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
LMArena WebDev14391639
LMArena Coding14601427
SWE-bench Verified75.4%—
FrontierCode—47.6%
SWE-bench Verified (bash only)75.8%—
CursorBench—46.3%
SWE-bench Multilingual72.7%—
FrontierSWE—29.5%
SciCode—57.8%
GSO9.8%—
WeirdML61.6%—
ALE-Bench1,367—

Agentic & Tool Use Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 38.7 (#29), Grok 4.7: 36.7 (#37)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
GDP.pdf10%22.8%
Vending-Bench 23,63510,537
Terminal-Bench64.3%—
APEX-Agents—54.6%
τ²-bench Airline82.5%—
τ²-bench Banking27.3%—
τ²-bench Retail76.8%—
τ²-bench Telecom91.2%—
DeepResearch Bench49.8%—
BALROG48.1%—
LMArena Search1198—

Reasoning Too close to call

Gemini 3 Flash Preview: 49.2 (#37), Grok 4.7: 49.1 (#40)

Reasoning benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
NYT Connections (extended)83.1%76.8%
Chess Puzzles40%38%
LMArena Hard Prompts14651413
Mystery Game Puzzles26%29%
DTBench89.1%96%
LMCA43.1%49.4%
Epoch Capabilities Index151.8153.53
ARC-AGI-233.6%—
SimpleBench61.1%—
ARC-AGI-184.7%—
CritPt—18%
ForecastBench58.5—

Math Grok 4.7 leads

Gemini 3 Flash Preview: 51.7 (#55), Grok 4.7: 57.8 (#39)

Math benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
FrontierMath (Tiers 1-3)51.2%53%
FrontierMath Tier 417.1%17.1%
OTIS Mock AIME 2024-202595.6%98.1%
ProofBench15%34%
LMArena Math14731407
MathArena Final-Answer Competitions67.6%—
FrontierMath (Feb 2025 set)35.6%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Grok 4.7 leads

Gemini 3 Flash Preview: 58.8 (#33), Grok 4.7: 62.8 (#22)

Knowledge benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
GPQA Diamond89.4%92.7%
SimpleQA Verified66.8%56%
LMArena Expert14621422
Vectara Hallucination Rate13.5%—

Multimodal Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 45.5 (#16), Grok 4.7: 35.5 (#87)

Multimodal benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
LMArena Vision12851228
Blueprint-Bench 20%32.5%
GeoBench88%—
VPCT72.6%—
Furniture Assembly—20.8%
LMArena Document1413—

Multilingual Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 55.7 (#27), Grok 4.7: 50.8 (#116)

Multilingual benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
LMArena Non-English14581389
LMArena Chinese15111455
LMArena French14771455
LMArena Russian14801397
LMArena Spanish14691400
LMArena German1497—
LMArena Japanese1489—
LMArena Korean1443—

Instruction Following Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 75.7 (#56), Grok 4.7: 74.1 (#105)

Instruction Following benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
LMArena Instruction Following14371404

Long Context Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 44.4 (#67), Grok 4.7: 43.1 (#104)

Long Context benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
LMArena Longer Query14521413

Writing & Preference Grok 4.7 leads

Gemini 3 Flash Preview: 65.5 (#45), Grok 4.7: 70.0 (#24)

Writing & Preference benchmarks
BenchmarkGemini 3 Flash PreviewGrok 4.7
LMArena Text14661399
LMArena Creative Writing14571391
LMArena Multi-Turn14711393
EQ-Bench Creative Writing—2007

Frequently asked questions

Is Gemini 3 Flash Preview better than Grok 4.7?

Gemini 3 Flash Preview and Grok 4.7 score almost the same on the Noometry Index (52.3 vs 53.1), so choose on price, context window or the category you care about most.

Which is cheaper, Gemini 3 Flash Preview or Grok 4.7?

Gemini 3 Flash Preview is cheaper. It lists at $0.50 per million input tokens and $3 per million output tokens; Grok 4.7 lists at $2 and $6.

Is Gemini 3 Flash Preview or Grok 4.7 better for coding?

Grok 4.7 scores higher on coding benchmarks: 58.0 versus 50.9 in the Noometry coding category.

Which has the bigger context window?

Gemini 3 Flash Preview does, with 1.05M tokens against 500K.

How many benchmarks do Gemini 3 Flash Preview and Grok 4.7 share?

31 benchmarks have published results for both models. Gemini 3 Flash Preview has 59 scored results on Noometry and Grok 4.7 has 39.

Related comparisons

Go deeper