Model comparison

Gemini 3.5 Flash vs GPT-5.6 Luna

Gemini 3.5 Flash and GPT-5.6 Luna score almost the same on the Noometry Index (54.2 vs 54.6), so choose on price, context window or the category you care about most.

Last verified . 46 shared benchmarks.

Gemini 3.5 Flash Google

54.2

Rank #32 Confirmed

GPT-5.6 Luna OpenAI

54.6

Rank #30 Confirmed

Summary

  • They share 46 benchmarks with published results for both. Gemini 3.5 Flash scores higher in 6 categories and GPT-5.6 Luna in 4 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.6 Luna leads 77.7 to 60.7.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 26.8% for Gemini 3.5 Flash and 61% for GPT-5.6 Luna.
  • GPT-5.6 Luna is cheaper at $0.20 / $1.20 per million input/output tokens, against $1.50 / $9 for Gemini 3.5 Flash.
  • GPT-5.6 Luna accepts more context: 1.05M tokens versus 1.05M.

Side by side

Gemini 3.5 Flash and GPT-5.6 Luna specifications
Gemini 3.5 FlashGPT-5.6 Luna
ProviderGoogleOpenAI
Noometry Index54.254.6
Released2026-05-192026-07-09
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output66K128K
Input $ / M tokens$1.50$0.20
Output $ / M tokens$9$1.20
Results tracked5452

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.6 Luna leads

Gemini 3.5 Flash: 49.4 (#49), GPT-5.6 Luna: 54.5 (#28)

Coding benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
DeepSWE37.4%67.2%
LMArena WebDev14991519
SciCode53.1%53.6%
WeirdML62.6%60.9%
LMArena Coding14921466
ALE-Bench911.021,667
SWE-bench Verified79.3%—
FrontierCode—39.8%
CursorBench—35.9%

Agentic & Tool Use GPT-5.6 Luna leads

Gemini 3.5 Flash: 24.7 (#114), GPT-5.6 Luna: 34.4 (#45)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
APEX-Agents27.5%43%
GDP.pdf14%22.7%
Vending-Bench 25,3964,095
BALROG—45.6%
GBAEval6.7%—

Reasoning Gemini 3.5 Flash leads

Gemini 3.5 Flash: 62.8 (#18), GPT-5.6 Luna: 47.6 (#43)

Reasoning benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
ARC-AGI-272.1%59.5%
SimpleBench76.7%46.8%
NYT Connections (extended)92.6%69.4%
ARC-AGI-192.5%88%
CritPt13.1%20.6%
Chess Puzzles50%40%
LMArena Hard Prompts14881451
Mystery Game Puzzles32%21%
DTBench94.7%89.1%
LMCA47.1%48.5%
Surface Evolver Bench58.1%61.9%
Epoch Capabilities Index154.46156.39
Kagi LLM Benchmark—49.1%
EnigmaEval25.4%—
EBR-Bench4.8%—
ForecastBench59—

Math GPT-5.6 Luna leads

Gemini 3.5 Flash: 60.7 (#36), GPT-5.6 Luna: 77.7 (#14)

Knowledge Gemini 3.5 Flash leads

Gemini 3.5 Flash: 66.3 (#11), GPT-5.6 Luna: 58.5 (#34)

Knowledge benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
GPQA Diamond92.8%91.6%
SimpleQA Verified66.2%41%
LMArena Expert14951478

Multimodal Gemini 3.5 Flash leads

Gemini 3.5 Flash: 45.7 (#15), GPT-5.6 Luna: 42.7 (#28)

Multimodal benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
LMArena Vision13101258
Blueprint-Bench 233.6%22.6%
LMArena Document14631457
Furniture Assembly—42.5%

Multilingual Gemini 3.5 Flash leads

Gemini 3.5 Flash: 57.0 (#13), GPT-5.6 Luna: 52.8 (#78)

Multilingual benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
LMArena Non-English14761417
LMArena Chinese15261470
LMArena French14901456
LMArena German14921454
LMArena Japanese14861411
LMArena Korean14511415
LMArena Russian14931428
LMArena Spanish14801448

Instruction Following Gemini 3.5 Flash leads

Gemini 3.5 Flash: 77.0 (#30), GPT-5.6 Luna: 75.6 (#57)

Instruction Following benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
LMArena Instruction Following14671437

Long Context Gemini 3.5 Flash leads

Gemini 3.5 Flash: 45.4 (#38), GPT-5.6 Luna: 43.9 (#82)

Long Context benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
LMArena Longer Query14821436

Writing & Preference GPT-5.6 Luna leads

Gemini 3.5 Flash: 65.5 (#47), GPT-5.6 Luna: 68.0 (#29)

Writing & Preference benchmarks
BenchmarkGemini 3.5 FlashGPT-5.6 Luna
LMArena Text14821431
LMArena Creative Writing14701396
EQ-Bench 410871156
LMArena Multi-Turn14811434
EQ-Bench Creative Writing—1829

Frequently asked questions

Is Gemini 3.5 Flash better than GPT-5.6 Luna?

Gemini 3.5 Flash and GPT-5.6 Luna score almost the same on the Noometry Index (54.2 vs 54.6), so choose on price, context window or the category you care about most.

Which is cheaper, Gemini 3.5 Flash or GPT-5.6 Luna?

GPT-5.6 Luna is cheaper. It lists at $0.20 per million input tokens and $1.20 per million output tokens; Gemini 3.5 Flash lists at $1.50 and $9.

Is Gemini 3.5 Flash or GPT-5.6 Luna better for coding?

GPT-5.6 Luna scores higher on coding benchmarks: 54.5 versus 49.4 in the Noometry coding category.

Which has the bigger context window?

GPT-5.6 Luna does, with 1.05M tokens against 1.05M.

How many benchmarks do Gemini 3.5 Flash and GPT-5.6 Luna share?

46 benchmarks have published results for both models. Gemini 3.5 Flash has 54 scored results on Noometry and GPT-5.6 Luna has 52.

Related comparisons

Go deeper