Model comparison

Gemini 3.1 Flash Lite vs GPT-5

GPT-5 is the stronger model overall, scoring 50.9 to 40.8 on the Noometry Index. Gemini 3.1 Flash Lite costs 6.1× less per token, which makes it the better buy when GPT-5's lead doesn't matter for your workload.

Last verified . 36 shared benchmarks.

Gemini 3.1 Flash Lite Google

40.8

Rank #144 Confirmed

GPT-5 OpenAI

50.9

Rank #45 Confirmed

Summary

  • They share 36 benchmarks with published results for both. Gemini 3.1 Flash Lite scores higher in 1 category and GPT-5 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in long context, where GPT-5 leads 69.5 to 42.5.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 27.7% for Gemini 3.1 Flash Lite and 55.4% for GPT-5.
  • Gemini 3.1 Flash Lite is cheaper at $0.25 / $1.50 per million input/output tokens, against $1.25 / $10 for GPT-5.
  • Gemini 3.1 Flash Lite accepts more context: 1.05M tokens versus 400K.

Side by side

Gemini 3.1 Flash Lite and GPT-5 specifications
Gemini 3.1 Flash LiteGPT-5
ProviderGoogleOpenAI
Noometry Index40.850.9
Released2026-03-032025-08-07
WeightsProprietaryProprietary
Context window1.05M400K
Max output66K128K
Input $ / M tokens$0.25$1.25
Output $ / M tokens$1.50$10
Results tracked3869

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5 leads

Gemini 3.1 Flash Lite: 37.8 (#188), GPT-5: 50.3 (#47)

Coding benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
LMArena WebDev12561418
SciCode41.9%42.9%
WeirdML52.2%60.7%
LMArena Coding14001436
ALE-Bench797.731,162
SWE-bench Verified—73.6%
SWE-bench Verified (bash only)—65%
Aider Polyglot—88%
GSO—6.9%
AlgoTune—1.67

Agentic & Tool Use GPT-5 leads

Gemini 3.1 Flash Lite: 30.2 (#79), GPT-5: 33.1 (#56)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
DeepResearch Bench37.3%49.6%
Terminal-Bench—49.6%
GDPval—34.8%
Remote Labor Index—1.7%
BALROG—32.8%
LMArena Search—1133
METR Time Horizons—69.6%

Reasoning GPT-5 leads

Gemini 3.1 Flash Lite: 22.9 (#186), GPT-5: 38.3 (#64)

Reasoning benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
Kagi LLM Benchmark67.2%72.7%
CritPt1.1%12.6%
Chess Puzzles25%37%
EnigmaEval3%10.5%
LMArena Hard Prompts14071416
DTBench76.8%90.7%
LMCA35%40%
Epoch Capabilities Index144.47150
ForecastBench54.461.4
ARC-AGI-2—9.9%
SimpleBench—56.7%
NYT Connections (extended)8.2%—
ARC-AGI-1—65.7%
Thematic Generalization63.3%—
EBR-Bench—12.7%
Mystery Game Puzzles—23%

Math GPT-5 leads

Gemini 3.1 Flash Lite: 40.7 (#90), GPT-5: 55.0 (#44)

Math benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
FrontierMath (Tiers 1-3)27.7%55.4%
OTIS Mock AIME 2024-202580%91.4%
LMArena Math14281407
FrontierMath Tier 4—22%
ProofBench—18%
Omni-MATH—64.7%
MATH Level 5—98.1%
FrontierMath (Feb 2025 set)—32.4%
FrontierMath Tier 4 (v1)—12.5%

Knowledge GPT-5 leads

Gemini 3.1 Flash Lite: 41.9 (#104), GPT-5: 56.6 (#43)

Knowledge benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
GPQA Diamond81.8%86.2%
Humanity's Last Exam8.6%25.3%
Vectara Hallucination Rate8.2%14.7%
LMArena Expert13981419
SimpleQA Verified—50.1%
MMLU-Pro—86.3%
Confabulations—10.3%
GPQA (HELM)—79.2%

Multimodal GPT-5 leads

Gemini 3.1 Flash Lite: 39.4 (#60), GPT-5: 46.8 (#13)

Multimodal benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
LMArena Vision12401232
GeoBench—81%
VPCT—66%

Multilingual Too close to call

Gemini 3.1 Flash Lite: 52.3 (#86), GPT-5: 51.4 (#110)

Multilingual benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
LMArena Non-English14111397
LMArena Chinese14611422
LMArena French14241410
LMArena German14291416
LMArena Japanese14131409
LMArena Korean13921360
LMArena Russian14201406
LMArena Spanish14211399

Instruction Following GPT-5 leads

Gemini 3.1 Flash Lite: 72.7 (#131), GPT-5: 73.8 (#113)

Instruction Following benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
LMArena Instruction Following13771388
IFEval—87.5%

Long Context GPT-5 leads

Gemini 3.1 Flash Lite: 42.5 (#122), GPT-5: 69.5 (#2)

Long Context benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
LMArena Longer Query13941399
Fiction.LiveBench—97.2%

Writing & Preference GPT-5 leads

Gemini 3.1 Flash Lite: 60.9 (#94), GPT-5: 63.4 (#65)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Flash LiteGPT-5
LMArena Text14161406
LMArena Creative Writing14011365
LMArena Multi-Turn14171426
Short-Story Creative Writing—86%
EQ-Bench Creative Writing—1627
WildBench—85.7%

Frequently asked questions

Is Gemini 3.1 Flash Lite better than GPT-5?

GPT-5 is the stronger model overall, scoring 50.9 to 40.8 on the Noometry Index. Gemini 3.1 Flash Lite costs 6.1× less per token, which makes it the better buy when GPT-5's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.1 Flash Lite or GPT-5?

Gemini 3.1 Flash Lite is cheaper. It lists at $0.25 per million input tokens and $1.50 per million output tokens; GPT-5 lists at $1.25 and $10.

Is Gemini 3.1 Flash Lite or GPT-5 better for coding?

GPT-5 scores higher on coding benchmarks: 50.3 versus 37.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 3.1 Flash Lite does, with 1.05M tokens against 400K.

How many benchmarks do Gemini 3.1 Flash Lite and GPT-5 share?

36 benchmarks have published results for both models. Gemini 3.1 Flash Lite has 38 scored results on Noometry and GPT-5 has 69.

Related comparisons

Go deeper