Model comparison

Gemini 3 Flash Preview vs GLM-5.2

Gemini 3 Flash Preview is the stronger model overall, scoring 52.3 to 51.1 on the Noometry Index.

Last verified . 39 shared benchmarks.

Gemini 3 Flash Preview Google

52.3

Rank #40 Confirmed

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Gemini 3 Flash Preview scores higher in 3 categories and GLM-5.2 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3 Flash Preview leads 49.2 to 42.3.
  • The biggest single-benchmark swing is SimpleQA Verified: 66.8% for Gemini 3 Flash Preview and 34.2% for GLM-5.2.
  • Gemini 3 Flash Preview is cheaper at $0.50 / $3 per million input/output tokens, against $1.40 / $4.40 for GLM-5.2.
  • Gemini 3 Flash Preview accepts more context: 1.05M tokens versus 1M.
  • GLM-5.2 has downloadable open weights; the other is API-only.

Side by side

Gemini 3 Flash Preview and GLM-5.2 specifications
Gemini 3 Flash PreviewGLM-5.2
ProviderGoogleZ.ai (Zhipu)
Noometry Index52.351.1
Released2025-12-172026-06-13
WeightsProprietaryOpen
Context window1.05M1M
Max output66K131K
Input $ / M tokens$0.50$1.40
Output $ / M tokens$3$4.40
Results tracked5951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 3 Flash Preview: 50.9 (#42), GLM-5.2: 51.3 (#41)

Coding benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
SWE-bench Verified75.4%78.7%
LMArena WebDev14391603
WeirdML61.6%70.1%
LMArena Coding14601485
ALE-Bench1,3671,047
DeepSWE—43.8%
FrontierCode—24.5%
SWE-bench Verified (bash only)75.8%—
SWE-bench Multilingual72.7%—
SciCode—50.5%
GSO9.8%—

Agentic & Tool Use Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 38.7 (#29), GLM-5.2: 32.4 (#63)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
τ²-bench Banking27.3%37.1%
Vending-Bench 23,6358,314
Terminal-Bench64.3%—
APEX-Agents—45.2%
τ²-bench Airline82.5%—
τ²-bench Retail76.8%—
τ²-bench Telecom91.2%—
DeepResearch Bench49.8%—
PostTrainBench—31.7%
BALROG48.1%—
GBAEval—0%
GDP.pdf10%—
LMArena Search1198—

Reasoning Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 49.2 (#37), GLM-5.2: 42.3 (#52)

Reasoning benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
ARC-AGI-233.6%22.8%
SimpleBench61.1%58.8%
NYT Connections (extended)83.1%74.3%
ARC-AGI-184.7%77%
Chess Puzzles40%21%
LMArena Hard Prompts14651480
Mystery Game Puzzles26%19%
DTBench89.1%93.6%
LMCA43.1%45.8%
Epoch Capabilities Index151.8151.78
Kagi LLM Benchmark—62.6%
CritPt—20.9%
EBR-Bench—9.5%
Surface Evolver Bench—55.6%
ForecastBench58.5—

Math GLM-5.2 leads

Gemini 3 Flash Preview: 51.7 (#55), GLM-5.2: 55.7 (#43)

Math benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
FrontierMath (Tiers 1-3)51.2%59.2%
FrontierMath Tier 417.1%29.3%
MathArena Final-Answer Competitions67.6%67.6%
OTIS Mock AIME 2024-202595.6%86.4%
ProofBench15%35%
LMArena Math14731482
FrontierMath (Feb 2025 set)35.6%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Gemini 3 Flash Preview leads

Gemini 3 Flash Preview: 58.8 (#33), GLM-5.2: 57.1 (#40)

Knowledge benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
GPQA Diamond89.4%91.9%
SimpleQA Verified66.8%34.2%
LMArena Expert14621486
Vectara Hallucination Rate13.5%—

Multimodal Not comparable

Gemini 3 Flash Preview: 45.5 (#16), GLM-5.2: —

Multimodal benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
LMArena Vision1285—
GeoBench88%—
VPCT72.6%—
Blueprint-Bench 20%—
LMArena Document1413—

Multilingual Too close to call

Gemini 3 Flash Preview: 55.7 (#27), GLM-5.2: 55.8 (#26)

Multilingual benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
LMArena Non-English14581459
LMArena Chinese15111519
LMArena French14771479
LMArena German14971468
LMArena Japanese14891451
LMArena Korean14431445
LMArena Russian14801466
LMArena Spanish14691477

Instruction Following GLM-5.2 leads

Gemini 3 Flash Preview: 75.7 (#56), GLM-5.2: 76.9 (#34)

Instruction Following benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
LMArena Instruction Following14371465

Long Context Too close to call

Gemini 3 Flash Preview: 44.4 (#67), GLM-5.2: 45.3 (#43)

Long Context benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
LMArena Longer Query14521479

Writing & Preference GLM-5.2 leads

Gemini 3 Flash Preview: 65.5 (#45), GLM-5.2: 70.4 (#21)

Writing & Preference benchmarks
BenchmarkGemini 3 Flash PreviewGLM-5.2
LMArena Text14661470
LMArena Creative Writing14571462
LMArena Multi-Turn14711469
EQ-Bench Creative Writing—1757
EQ-Bench 4—1222

Frequently asked questions

Is Gemini 3 Flash Preview better than GLM-5.2?

Gemini 3 Flash Preview is the stronger model overall, scoring 52.3 to 51.1 on the Noometry Index.

Which is cheaper, Gemini 3 Flash Preview or GLM-5.2?

Gemini 3 Flash Preview is cheaper. It lists at $0.50 per million input tokens and $3 per million output tokens; GLM-5.2 lists at $1.40 and $4.40.

Is Gemini 3 Flash Preview or GLM-5.2 better for coding?

They score almost the same on coding (50.9 vs 51.3); test both on your own repository before choosing.

Which has the bigger context window?

Gemini 3 Flash Preview does, with 1.05M tokens against 1M.

How many benchmarks do Gemini 3 Flash Preview and GLM-5.2 share?

39 benchmarks have published results for both models. Gemini 3 Flash Preview has 59 scored results on Noometry and GLM-5.2 has 51.

Related comparisons

Go deeper