Model comparison

Gemini 3.1 Pro Preview vs Mistral Large

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 31.9 on the Noometry Index. Mistral Large costs 1.5× less per token, which makes it the better buy when Gemini 3.1 Pro Preview's lead doesn't matter for your workload.

Last verified . 30 shared benchmarks.

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Gemini 3.1 Pro Preview scores higher in 9 categories and Mistral Large in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.1 Pro Preview leads 71.7 to 15.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 95.6% for Gemini 3.1 Pro Preview and 8.5% for Mistral Large.
  • Mistral Large is cheaper at $2 / $6 per million input/output tokens, against $2 / $12 for Gemini 3.1 Pro Preview.
  • Gemini 3.1 Pro Preview accepts more context: 1.05M tokens versus 131K.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Gemini 3.1 Pro Preview and Mistral Large specifications
Gemini 3.1 Pro PreviewMistral Large
ProviderGoogleMistral AI
Noometry Index56.731.9
Released2026-02-192024-02-26
WeightsProprietaryOpen
Context window1.05M131K
Max output66K16K
Input $ / M tokens$2$2
Output $ / M tokens$12$6
Results tracked7151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 42.5 (#99), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
SciCode58.9%36.2%
LMArena Coding14841277
ALE-Bench1,161264.7
SWE-bench Verified75.6%—
DeepSWE11.7%—
LMArena WebDev1447—
GSO22.6%—
WeirdML72.1%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
MirrorCode8.9%—
BigCodeBench Complete—38.3%
AlgoTune2.02—
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 37.7 (#34), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
Terminal-Bench80.2%—
APEX-Agents35.3%—
Berkeley Function Calling Leaderboard—38.4%
τ²-bench Banking26%—
DeepResearch Bench47.8%—
PostTrainBench22%—
BALROG57%—
ExploitBench26.1%—
GBAEval0.8%—
GDP.pdf17%—
LMArena Search1211—
METR Time Horizons77%—
Vending-Bench 23,774—

Reasoning Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.7 (#12), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
SimpleBench79.6%22.5%
CritPt17.7%0%
LMArena Hard Prompts14851257
DTBench97.1%65.1%
LMCA53.8%16.7%
Epoch Capabilities Index154.77128.52
ForecastBench5957.1
ARC-AGI-277.1%—
NYT Connections (extended)97.4%—
ARC-AGI-198%—
Chess Puzzles55%—
EnigmaEval36.8%—
Thematic Generalization79.4%—
EBR-Bench14.3%—
LiveBench Reasoning—43.5%
Mystery Game Puzzles34%—
LiveBench Data Analysis—50.1%
LiveBench—48.4%

Math Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 62.1 (#34), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
OTIS Mock AIME 2024-202595.6%8.5%
LMArena Math14851262
FrontierMath (Feb 2025 set)36.9%0.3%
FrontierMath (Tiers 1-3)59.6%—
FrontierMath Tier 426.8%—
MathArena Final-Answer Competitions86.5%—
ProofBench26%—
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath Tier 4 (v1)16.7%—

Knowledge Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.8 (#3), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
GPQA Diamond94.4%51.3%
Vectara Hallucination Rate10.4%4.5%
LMArena Expert14851232
Humanity's Last Exam46.4%—
SimpleQA Verified73.5%—
MMLU-Pro—59.9%
Confabulations—21.4%
GPQA (HELM)—43.5%
MMLU—80%

Multimodal Not comparable

Gemini 3.1 Pro Preview: 37.9 (#69), Mistral Large: —

Multimodal benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
LMArena Vision1296—
Blueprint-Bench 226.5%—
Furniture Assembly26.7%—
LMArena Document1444—

Multilingual Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 57.0 (#12), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
LMArena Non-English14771237
LMArena Chinese15291240
LMArena French14871325
LMArena German14911254
LMArena Japanese14931188
LMArena Korean14551202
LMArena Russian14981257
LMArena Spanish14791268

Instruction Following Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 77.0 (#32), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
LMArena Instruction Following14661249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 47.4 (#18), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
LMArena Longer Query14831261
CL-bench20.8%—
CL-bench Life16.9%—

Writing & Preference Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 66.1 (#37), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Pro PreviewMistral Large
LMArena Text14811266
LMArena Creative Writing14821243
EQ-Bench Creative Writing1491985
LMArena Multi-Turn14881260
Short-Story Creative Writing—69%
WildBench—80.1%
EQ-Bench 41142—
LiveBench Language—39.4%

Frequently asked questions

Is Gemini 3.1 Pro Preview better than Mistral Large?

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 31.9 on the Noometry Index. Mistral Large costs 1.5× less per token, which makes it the better buy when Gemini 3.1 Pro Preview's lead doesn't matter for your workload.

Which is cheaper, Gemini 3.1 Pro Preview or Mistral Large?

Mistral Large is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Gemini 3.1 Pro Preview lists at $2 and $12.

Is Gemini 3.1 Pro Preview or Mistral Large better for coding?

Gemini 3.1 Pro Preview scores higher on coding benchmarks: 42.5 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Gemini 3.1 Pro Preview does, with 1.05M tokens against 131K.

How many benchmarks do Gemini 3.1 Pro Preview and Mistral Large share?

30 benchmarks have published results for both models. Gemini 3.1 Pro Preview has 71 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper