Model comparison

Chatgpt 4o Latest 20250326 vs Gemini 3.5 Flash Lite

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 41.5 on the Noometry Index.

Last verified . 19 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

Gemini 3.5 Flash Lite Google

41.5

Rank #133 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 2 categories and Gemini 3.5 Flash Lite in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Chatgpt 4o Latest 20250326 leads 38.6 to 25.9.

Side by side

Chatgpt 4o Latest 20250326 and Gemini 3.5 Flash Lite specifications
Chatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
ProviderOpenAIGoogle
Noometry Index43.841.5
Released—2026-07-21
WeightsProprietaryProprietary
Context window—1.05M
Max output—66K
Input $ / M tokens—$0.30
Output $ / M tokens—$2.50
Results tracked2139

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Chatgpt 4o Latest 20250326: 41.6 (#122), Gemini 3.5 Flash Lite: 42.4 (#104)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Coding14131453
LMArena WebDev—1440
SciCode—41.3%
WeirdML—39%
ALE-Bench—765.27

Agentic & Tool Use Not comparable

Chatgpt 4o Latest 20250326: —, Gemini 3.5 Flash Lite: 25.7 (#107)

Agentic & Tool Use benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
APEX-Agents—29.3%
GDP.pdf—10%

Reasoning Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 33.8 (#71), Gemini 3.5 Flash Lite: 27.8 (#114)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Hard Prompts14241441
ARC-AGI-2—10.3%
Kagi LLM Benchmark75%—
NYT Connections (extended)—60.4%
ARC-AGI-1—53.5%
CritPt—0%
Chess Puzzles—22%
Mystery Game Puzzles—19%
DTBench—83.5%
LMCA—37.6%
Epoch Capabilities Index—145.13

Math Chatgpt 4o Latest 20250326 leads

Chatgpt 4o Latest 20250326: 38.6 (#134), Gemini 3.5 Flash Lite: 25.9 (#264)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Math14071433
FrontierMath (Tiers 1-3)—26%
FrontierMath Tier 4—0%
OTIS Mock AIME 2024-2025—71.1%
ProofBench—13%

Knowledge Gemini 3.5 Flash Lite leads

Chatgpt 4o Latest 20250326: 39.2 (#137), Gemini 3.5 Flash Lite: 50.1 (#72)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Expert14011435
GPQA Diamond—83.3%
Confabulations16.6%—

Multimodal Gemini 3.5 Flash Lite leads

Chatgpt 4o Latest 20250326: 39.6 (#58), Gemini 3.5 Flash Lite: 41.1 (#40)

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Vision12431268

Multilingual Too close to call

Chatgpt 4o Latest 20250326: 52.9 (#76), Gemini 3.5 Flash Lite: 53.5 (#63)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Non-English14191427
LMArena Chinese14571469
LMArena French14461451
LMArena German14231454
LMArena Japanese14051427
LMArena Korean13961407
LMArena Russian14291443
LMArena Spanish14341442

Instruction Following Too close to call

Chatgpt 4o Latest 20250326: 74.0 (#107), Gemini 3.5 Flash Lite: 74.8 (#85)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Instruction Following14031420

Long Context Too close to call

Chatgpt 4o Latest 20250326: 43.1 (#107), Gemini 3.5 Flash Lite: 43.9 (#84)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Longer Query14131435

Writing & Preference Gemini 3.5 Flash Lite leads

Chatgpt 4o Latest 20250326: 62.6 (#72), Gemini 3.5 Flash Lite: 64.2 (#57)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326Gemini 3.5 Flash Lite
LMArena Text14291435
LMArena Creative Writing14051420
EQ-Bench Creative Writing15011559
LMArena Multi-Turn14541445

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than Gemini 3.5 Flash Lite?

Chatgpt 4o Latest 20250326 is the stronger model overall, scoring 43.8 to 41.5 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or Gemini 3.5 Flash Lite better for coding?

They score almost the same on coding (41.6 vs 42.4); test both on your own repository before choosing.

How many benchmarks do Chatgpt 4o Latest 20250326 and Gemini 3.5 Flash Lite share?

19 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and Gemini 3.5 Flash Lite has 39.

Related comparisons

Go deeper