Model comparison

Gemini 3.1 Pro Preview vs GPT-5.6 Sol

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 56.7 on the Noometry Index. Gemini 3.1 Pro Preview costs 1.8× less per token, which makes it the better buy when GPT-5.6 Sol's lead doesn't matter for your workload.

Last verified . 57 shared benchmarks.

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

GPT-5.6 Sol OpenAI

65.0

Rank #7 Confirmed

Summary

  • They share 57 benchmarks with published results for both. Gemini 3.1 Pro Preview scores higher in 3 categories and GPT-5.6 Sol in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.6 Sol leads 85.6 to 62.1.
  • The biggest single-benchmark swing is DeepSWE: 11.7% for Gemini 3.1 Pro Preview and 72.7% for GPT-5.6 Sol.
  • Gemini 3.1 Pro Preview is cheaper at $2 / $12 per million input/output tokens, against $4 / $20 for GPT-5.6 Sol.
  • GPT-5.6 Sol accepts more context: 1.05M tokens versus 1.05M.

Side by side

Gemini 3.1 Pro Preview and GPT-5.6 Sol specifications
Gemini 3.1 Pro PreviewGPT-5.6 Sol
ProviderGoogleOpenAI
Noometry Index56.765.0
Released2026-02-192026-07-09
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output66K128K
Input $ / M tokens$2$4
Output $ / M tokens$12$20
Results tracked7165

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.6 Sol leads

Gemini 3.1 Pro Preview: 42.5 (#99), GPT-5.6 Sol: 65.1 (#7)

Coding benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
DeepSWE11.7%72.7%
LMArena WebDev14471618
SciCode58.9%57.1%
GSO22.6%76.5%
WeirdML72.1%89.4%
LMArena Coding14841498
MirrorCode8.9%20%
ALE-Bench1,1612,177
SWE-bench Verified75.6%—
FrontierCode—47.5%
CursorBench—41.7%
FrontierSWE—32.2%
AlgoTune2.02—

Agentic & Tool Use GPT-5.6 Sol leads

Gemini 3.1 Pro Preview: 37.7 (#34), GPT-5.6 Sol: 50.3 (#7)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
APEX-Agents35.3%51.4%
τ²-bench Banking26%46.9%
PostTrainBench22%36.2%
BALROG57%60%
GBAEval0.8%52.6%
GDP.pdf17%30.7%
LMArena Search12111257
Vending-Bench 23,7749,619
Terminal-Bench80.2%—
OSWorld 2.0—27.3%
DeepResearch Bench47.8%—
ExploitBench26.1%—
METR Time Horizons77%—

Reasoning GPT-5.6 Sol leads

Gemini 3.1 Pro Preview: 71.7 (#12), GPT-5.6 Sol: 74.8 (#8)

Reasoning benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
ARC-AGI-277.1%92.5%
SimpleBench79.6%71.7%
NYT Connections (extended)97.4%93.8%
ARC-AGI-198%97.5%
CritPt17.7%32.3%
Chess Puzzles55%64%
EnigmaEval36.8%37.1%
EBR-Bench14.3%44.8%
LMArena Hard Prompts14851484
Mystery Game Puzzles34%58%
DTBench97.1%96%
LMCA53.8%59.2%
Epoch Capabilities Index154.77161.66
Kagi LLM Benchmark—67%
Thematic Generalization79.4%—
Surface Evolver Bench—93.1%
Bench to the Future 3—0.14
ForecastBench59—

Math GPT-5.6 Sol leads

Gemini 3.1 Pro Preview: 62.1 (#34), GPT-5.6 Sol: 85.6 (#9)

Knowledge Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.8 (#3), GPT-5.6 Sol: 64.3 (#18)

Knowledge benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
GPQA Diamond94.4%93.5%
SimpleQA Verified73.5%69.7%
Vectara Hallucination Rate10.4%12.4%
LMArena Expert14851516
Humanity's Last Exam46.4%—

Multimodal GPT-5.6 Sol leads

Gemini 3.1 Pro Preview: 37.9 (#69), GPT-5.6 Sol: 48.6 (#9)

Multimodal benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
LMArena Vision12961281
Blueprint-Bench 226.5%33.6%
Furniture Assembly26.7%56.7%
LMArena Document14441483

Multilingual Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 57.0 (#12), GPT-5.6 Sol: 55.3 (#32)

Multilingual benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
LMArena Non-English14771452
LMArena Chinese15291527
LMArena French14871477
LMArena German14911476
LMArena Japanese14931471
LMArena Korean14551442
LMArena Russian14981468
LMArena Spanish14791441

Instruction Following Too close to call

Gemini 3.1 Pro Preview: 77.0 (#32), GPT-5.6 Sol: 77.7 (#16)

Instruction Following benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
LMArena Instruction Following14661482

Long Context Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 47.4 (#18), GPT-5.6 Sol: 45.4 (#42)

Long Context benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
LMArena Longer Query14831480
CL-bench20.8%—
CL-bench Life16.9%—

Writing & Preference GPT-5.6 Sol leads

Gemini 3.1 Pro Preview: 66.1 (#37), GPT-5.6 Sol: 73.3 (#12)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Pro PreviewGPT-5.6 Sol
LMArena Text14811457
LMArena Creative Writing14821448
EQ-Bench Creative Writing14911972
EQ-Bench 411421250
LMArena Multi-Turn14881460

Frequently asked questions

Is Gemini 3.1 Pro Preview better than GPT-5.6 Sol?

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 56.7 on the Noometry Index. Gemini 3.1 Pro Preview costs 1.8× less per token, which makes it the better buy when GPT-5.6 Sol's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.1 Pro Preview or GPT-5.6 Sol?

Gemini 3.1 Pro Preview is cheaper. It lists at $2 per million input tokens and $12 per million output tokens; GPT-5.6 Sol lists at $4 and $20.

Is Gemini 3.1 Pro Preview or GPT-5.6 Sol better for coding?

GPT-5.6 Sol scores higher on coding benchmarks: 65.1 versus 42.5 in the Noometry coding category.

Which has the bigger context window?

GPT-5.6 Sol does, with 1.05M tokens against 1.05M.

How many benchmarks do Gemini 3.1 Pro Preview and GPT-5.6 Sol share?

57 benchmarks have published results for both models. Gemini 3.1 Pro Preview has 71 scored results on Noometry and GPT-5.6 Sol has 65.

Related comparisons

Go deeper