Model comparison

Gemini 3.5 Flash vs GPT-5.5 Instant

Gemini 3.5 Flash is the stronger model overall, scoring 54.2 to 42.7 on the Noometry Index.

Last verified . 27 shared benchmarks.

Gemini 3.5 Flash Google

54.2

Rank #32 Confirmed

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Gemini 3.5 Flash scores higher in 9 categories and GPT-5.5 Instant in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.5 Flash leads 62.8 to 24.9.
  • The biggest single-benchmark swing is Chess Puzzles: 50% for Gemini 3.5 Flash and 12% for GPT-5.5 Instant.

Side by side

Gemini 3.5 Flash and GPT-5.5 Instant specifications
Gemini 3.5 FlashGPT-5.5 Instant
ProviderGoogleOpenAI
Noometry Index54.242.7
Released2026-05-192026-05-05
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$1.50—
Output $ / M tokens$9—
Results tracked5427

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3.5 Flash leads

Gemini 3.5 Flash: 49.4 (#49), GPT-5.5 Instant: 44.3 (#74)

Coding benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
SciCode53.1%48.6%
LMArena Coding14921433
SWE-bench Verified79.3%—
DeepSWE37.4%—
LMArena WebDev1499—
WeirdML62.6%—
ALE-Bench911.02—

Agentic & Tool Use Not comparable

Gemini 3.5 Flash: 24.7 (#114), GPT-5.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
APEX-Agents27.5%—
GBAEval6.7%—
GDP.pdf14%—
Vending-Bench 25,396—

Reasoning Gemini 3.5 Flash leads

Gemini 3.5 Flash: 62.8 (#18), GPT-5.5 Instant: 24.9 (#155)

Reasoning benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
CritPt13.1%0%
Chess Puzzles50%12%
LMArena Hard Prompts14881426
Epoch Capabilities Index154.46142.52
ARC-AGI-272.1%—
SimpleBench76.7%—
NYT Connections (extended)92.6%—
ARC-AGI-192.5%—
EnigmaEval25.4%—
EBR-Bench4.8%—
Mystery Game Puzzles32%—
DTBench94.7%—
LMCA47.1%—
Surface Evolver Bench58.1%—
ForecastBench59—

Math Gemini 3.5 Flash leads

Gemini 3.5 Flash: 60.7 (#36), GPT-5.5 Instant: 26.5 (#259)

Math benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
FrontierMath (Tiers 1-3)62.8%26.3%
FrontierMath Tier 426.8%2.4%
OTIS Mock AIME 2024-202595.6%68.1%
LMArena Math15041420
MathArena Final-Answer Competitions76.3%—
ProofBench31%—
FrontierMath (Feb 2025 set)39%—
FrontierMath Tier 4 (v1)14.6%—

Knowledge Gemini 3.5 Flash leads

Gemini 3.5 Flash: 66.3 (#11), GPT-5.5 Instant: 48.9 (#74)

Knowledge benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
GPQA Diamond92.8%82.5%
LMArena Expert14951409
SimpleQA Verified66.2%—

Multimodal Gemini 3.5 Flash leads

Gemini 3.5 Flash: 45.7 (#15), GPT-5.5 Instant: 40.0 (#52)

Multimodal benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
LMArena Vision13101250
LMArena Document14631403
Blueprint-Bench 233.6%—

Multilingual Gemini 3.5 Flash leads

Gemini 3.5 Flash: 57.0 (#13), GPT-5.5 Instant: 52.8 (#80)

Multilingual benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
LMArena Non-English14761417
LMArena Chinese15261456
LMArena French14901428
LMArena German14921411
LMArena Japanese14861408
LMArena Korean14511392
LMArena Russian14931431
LMArena Spanish14801429

Instruction Following Gemini 3.5 Flash leads

Gemini 3.5 Flash: 77.0 (#30), GPT-5.5 Instant: 74.2 (#100)

Instruction Following benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
LMArena Instruction Following14671406

Long Context Gemini 3.5 Flash leads

Gemini 3.5 Flash: 45.4 (#38), GPT-5.5 Instant: 43.4 (#96)

Long Context benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
LMArena Longer Query14821422

Writing & Preference Gemini 3.5 Flash leads

Gemini 3.5 Flash: 65.5 (#47), GPT-5.5 Instant: 61.8 (#85)

Writing & Preference benchmarks
BenchmarkGemini 3.5 FlashGPT-5.5 Instant
LMArena Text14821419
LMArena Creative Writing14701419
LMArena Multi-Turn14811433
EQ-Bench 41087—

Frequently asked questions

Is Gemini 3.5 Flash better than GPT-5.5 Instant?

Gemini 3.5 Flash is the stronger model overall, scoring 54.2 to 42.7 on the Noometry Index.

Is Gemini 3.5 Flash or GPT-5.5 Instant better for coding?

Gemini 3.5 Flash scores higher on coding benchmarks: 49.4 versus 44.3 in the Noometry coding category.

How many benchmarks do Gemini 3.5 Flash and GPT-5.5 Instant share?

27 benchmarks have published results for both models. Gemini 3.5 Flash has 54 scored results on Noometry and GPT-5.5 Instant has 27.

Related comparisons

Go deeper