Model comparison

Gemini 3.8 Flash vs Longcat Flash Chat

Gemini 3.8 Flash is the stronger model overall, scoring 61.8 to 42.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

Gemini 3.8 Flash Google

61.8

Rank #11 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemini 3.8 Flash scores higher in 8 categories and Longcat Flash Chat in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.8 Flash leads 76.9 to 19.0.
  • The biggest single-benchmark swing is NYT Connections (extended): 97.4% for Gemini 3.8 Flash and 17.7% for Longcat Flash Chat.
  • Longcat Flash Chat has downloadable open weights; the other is API-only.

Side by side

Gemini 3.8 Flash and Longcat Flash Chat specifications
Gemini 3.8 FlashLongcat Flash Chat
ProviderGoogleMeituan
Noometry Index61.842.1
Released2026-09-02—
WeightsProprietaryOpen
Context window1.05M—
Max output66K—
Input $ / M tokens$0.75—
Output $ / M tokens$3.75—
Results tracked5019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3.8 Flash leads

Gemini 3.8 Flash: 59.2 (#15), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Coding15101471
DeepSWE73.8%—
FrontierCode41.2%—
CursorBench39.6%—
LMArena WebDev1584—
FrontierSWE19.6%—
SciCode56.6%—
WeirdML84.8%—
ALE-Bench1,270—

Agentic & Tool Use Not comparable

Gemini 3.8 Flash: 41.8 (#21), Longcat Flash Chat: —

Agentic & Tool Use benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
APEX-Agents64.3%—
Remote Labor Index5.8%—
GDP.pdf23.4%—
Vending-Bench 25,094—

Reasoning Gemini 3.8 Flash leads

Gemini 3.8 Flash: 76.9 (#5), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
NYT Connections (extended)97.4%17.7%
LMArena Hard Prompts15081440
ARC-AGI-289.2%—
Kagi LLM Benchmark—43.9%
ARC-AGI-198.5%—
CritPt18.3%—
Chess Puzzles61%—
Mystery Game Puzzles47%—
DTBench95.7%—
LMCA52.9%—
Surface Evolver Bench76.9%—
Epoch Capabilities Index156.71—

Math Gemini 3.8 Flash leads

Gemini 3.8 Flash: 65.3 (#28), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Math15281442
FrontierMath (Tiers 1-3)68.4%—
FrontierMath Tier 422%—
OTIS Mock AIME 2024-202598.9%—
ProofBench48%—

Knowledge Gemini 3.8 Flash leads

Gemini 3.8 Flash: 74.8 (#2), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Expert15241454
GPQA Diamond95.4%—
Humanity's Last Exam44.5%—
SimpleQA Verified69.7%—

Multimodal Not comparable

Gemini 3.8 Flash: 40.7 (#45), Longcat Flash Chat: —

Multimodal benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Vision1314—
Blueprint-Bench 238.6%—
Furniture Assembly31.7%—

Multilingual Gemini 3.8 Flash leads

Gemini 3.8 Flash: 58.0 (#5), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Non-English14911404
LMArena Chinese15541465
LMArena French14981456
LMArena German14931408
LMArena Japanese15021373
LMArena Korean14591371
LMArena Russian15151395
LMArena Spanish14851445

Instruction Following Gemini 3.8 Flash leads

Gemini 3.8 Flash: 78.0 (#13), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Instruction Following14901411

Long Context Gemini 3.8 Flash leads

Gemini 3.8 Flash: 46.3 (#24), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Longer Query15081425

Writing & Preference Gemini 3.8 Flash leads

Gemini 3.8 Flash: 72.2 (#15), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkGemini 3.8 FlashLongcat Flash Chat
LMArena Text14991427
LMArena Creative Writing14921388
LMArena Multi-Turn15011418
EQ-Bench Creative Writing1748—

Frequently asked questions

Is Gemini 3.8 Flash better than Longcat Flash Chat?

Gemini 3.8 Flash is the stronger model overall, scoring 61.8 to 42.1 on the Noometry Index.

Is Gemini 3.8 Flash or Longcat Flash Chat better for coding?

Gemini 3.8 Flash scores higher on coding benchmarks: 59.2 versus 43.5 in the Noometry coding category.

How many benchmarks do Gemini 3.8 Flash and Longcat Flash Chat share?

18 benchmarks have published results for both models. Gemini 3.8 Flash has 50 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper