Model comparison

Gemini 3 Flash Preview vs Kimi K2.6

Gemini 3 Flash Preview is the stronger model overall, scoring 52.3 to 47.7 on the Noometry Index.

Last verified . 42 shared benchmarks.

Gemini 3 Flash Preview Google

52.3

Rank #40 Confirmed

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Summary

  • They share 42 benchmarks with published results for both. Gemini 3 Flash Preview scores higher in 6 categories and Kimi K2.6 in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Gemini 3 Flash Preview leads 38.7 to 21.9.
  • The biggest single-benchmark swing is SimpleQA Verified: 66.8% for Gemini 3 Flash Preview and 34.9% for Kimi K2.6.
  • Gemini 3 Flash Preview is cheaper at $0.50 / $3 per million input/output tokens, against $0.95 / $4 for Kimi K2.6.
  • Gemini 3 Flash Preview accepts more context: 1.05M tokens versus 262K.
  • Kimi K2.6 has downloadable open weights; the other is API-only.

Side by side

Gemini 3 Flash Preview and Kimi K2.6 specifications
Gemini 3 Flash PreviewKimi K2.6
ProviderGoogleMoonshot AI
Noometry Index52.347.7
Released2025-12-172026-04-20
WeightsProprietaryOpen
Context window1.05M262K
Max output66K262K
Input $ / M tokens$0.50$0.95
Output $ / M tokens$3$4
Results tracked5951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 3 Flash Preview: 50.9 (#42), Kimi K2.6: 50.7 (#43)

Coding benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
SWE-bench Verified75.4%76.7%
LMArena WebDev14391509
WeirdML61.6%55.9%
LMArena Coding14601488
ALE-Bench1,3671,093
SWE-bench Verified (bash only)75.8%—
SWE-bench Multilingual72.7%—
SciCode—53.5%
GSO9.8%—

Agentic & Tool Use Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 38.7 (#29), Kimi K2.6: 21.9 (#137)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
GDP.pdf10%12%
Vending-Bench 23,6356,205
Terminal-Bench64.3%—
OSWorld 2.0—4.6%
τ²-bench Airline82.5%—
τ²-bench Banking27.3%—
τ²-bench Retail76.8%—
τ²-bench Telecom91.2%—
DeepResearch Bench49.8%—
BALROG48.1%—
ExploitBench—18.4%
GBAEval—0.9%
LMArena Search1198—

Reasoning Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 49.2 (#37), Kimi K2.6: 40.5 (#55)

Reasoning benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
NYT Connections (extended)83.1%87.2%
Chess Puzzles40%26%
LMArena Hard Prompts14651470
Mystery Game Puzzles26%18%
DTBench89.1%90.9%
LMCA43.1%37.3%
Epoch Capabilities Index151.8151.05
ARC-AGI-233.6%—
SimpleBench61.1%—
ARC-AGI-184.7%—
CritPt—8%
EBR-Bench—2.4%
ForecastBench58.5—

Math Kimi K2.6 leads

Gemini 3 Flash Preview: 51.7 (#55), Kimi K2.6: 57.0 (#41)

Math benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
FrontierMath (Tiers 1-3)51.2%57.2%
FrontierMath Tier 417.1%25.6%
MathArena Final-Answer Competitions67.6%72.9%
OTIS Mock AIME 2024-202595.6%96.1%
ProofBench15%16%
LMArena Math14731475
FrontierMath (Feb 2025 set)35.6%39%
FrontierMath Tier 4 (v1)4.2%14.6%

Knowledge Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 58.8 (#33), Kimi K2.6: 54.0 (#54)

Knowledge benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
GPQA Diamond89.4%90.8%
SimpleQA Verified66.8%34.9%
Vectara Hallucination Rate13.5%10.8%
LMArena Expert14621491

Multimodal Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 45.5 (#16), Kimi K2.6: 31.6 (#103)

Multimodal benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
LMArena Vision12851283
Blueprint-Bench 20%3.9%
LMArena Document14131451
GeoBench88%—
VPCT72.6%—
Furniture Assembly—21.7%

Multilingual Too close to call

Gemini 3 Flash Preview: 55.7 (#27), Kimi K2.6: 54.9 (#37)

Multilingual benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
LMArena Non-English14581446
LMArena Chinese15111521
LMArena French14771471
LMArena German14971450
LMArena Japanese14891443
LMArena Korean14431427
LMArena Russian14801446
LMArena Spanish14691464

Instruction Following Too close to call

Gemini 3 Flash Preview: 75.7 (#56), Kimi K2.6: 76.3 (#43)

Instruction Following benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
LMArena Instruction Following14371451

Long Context Too close to call

Gemini 3 Flash Preview: 44.4 (#67), Kimi K2.6: 44.9 (#52)

Long Context benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
LMArena Longer Query14521468

Writing & Preference Kimi K2.6 leads

Gemini 3 Flash Preview: 65.5 (#45), Kimi K2.6: 68.5 (#26)

Writing & Preference benchmarks
BenchmarkGemini 3 Flash PreviewKimi K2.6
LMArena Text14661455
LMArena Creative Writing14571434
LMArena Multi-Turn14711453
EQ-Bench Creative Writing—1725
EQ-Bench 4—1202

Frequently asked questions

Is Gemini 3 Flash Preview better than Kimi K2.6?

Gemini 3 Flash Preview is the stronger model overall, scoring 52.3 to 47.7 on the Noometry Index.

Which is cheaper, Gemini 3 Flash Preview or Kimi K2.6?

Gemini 3 Flash Preview is cheaper. It lists at $0.50 per million input tokens and $3 per million output tokens; Kimi K2.6 lists at $0.95 and $4.

Is Gemini 3 Flash Preview or Kimi K2.6 better for coding?

They score almost the same on coding (50.9 vs 50.7); test both on your own repository before choosing.

Which has the bigger context window?

Gemini 3 Flash Preview does, with 1.05M tokens against 262K.

How many benchmarks do Gemini 3 Flash Preview and Kimi K2.6 share?

42 benchmarks have published results for both models. Gemini 3 Flash Preview has 59 scored results on Noometry and Kimi K2.6 has 51.

Related comparisons

Go deeper