Model comparison

Gemma 3 12B vs Mistral Large

Gemma 3 12B and Mistral Large score almost the same on the Noometry Index (32.1 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

Gemma 3 12B Google

32.1

Rank #262 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Gemma 3 12B scores higher in 5 categories and Mistral Large in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Gemma 3 12B leads 47.5 to 40.7.
  • The biggest single-benchmark swing is SciCode: 17.4% for Gemma 3 12B and 36.2% for Mistral Large.
  • Gemma 3 12B is cheaper at $0.05 / $0.15 per million input/output tokens, against $2 / $6 for Mistral Large.

Side by side

Gemma 3 12B and Mistral Large specifications
Gemma 3 12BMistral Large
ProviderGoogleMistral AI
Noometry Index32.131.9
Released2025-03-122024-02-26
WeightsOpenOpen
Context window131K131K
Max output8K16K
Input $ / M tokens$0.05$2
Output $ / M tokens$0.15$6
Results tracked2451

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Gemma 3 12B: 31.7 (#280), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkGemma 3 12BMistral Large
SciCode17.4%36.2%
LMArena Coding12811277
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

Gemma 3 12B: 25.5 (#108), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 12BMistral Large
Berkeley Function Calling Leaderboard30.4%38.4%

Reasoning Too close to call

Gemma 3 12B: 15.7 (#313), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkGemma 3 12BMistral Large
CritPt0%0%
LMArena Hard Prompts13091257
DTBench48.8%65.1%
LMCA4.5%16.7%
Epoch Capabilities Index123.5128.52
SimpleBench—22.5%
Chess Puzzles0%—
LiveBench Reasoning—43.5%
LiveBench Data Analysis—50.1%
ForecastBench—57.1
LiveBench—48.4%

Math Gemma 3 12B leads

Gemma 3 12B: 22.3 (#279), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkGemma 3 12BMistral Large
OTIS Mock AIME 2024-202516.7%8.5%
LMArena Math13071262
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Mistral Large leads

Gemma 3 12B: 26.5 (#257), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkGemma 3 12BMistral Large
GPQA Diamond39.5%51.3%
Vectara Hallucination Rate4.4%4.5%
LMArena Expert12481232
MMLU-Pro—59.9%
Confabulations—21.4%
GPQA (HELM)—43.5%
MMLU—80%

Multimodal Not comparable

Gemma 3 12B: —, Mistral Large: —

Multimodal benchmarks
BenchmarkGemma 3 12BMistral Large
MindCube46.7%—

Multilingual Gemma 3 12B leads

Gemma 3 12B: 45.7 (#165), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkGemma 3 12BMistral Large
LMArena Non-English13181237
LMArena German13701254
LMArena Russian13351257
LMArena Chinese—1240
LMArena French—1325
LMArena Japanese—1188
LMArena Korean—1202
LMArena Spanish—1268

Instruction Following Too close to call

Gemma 3 12B: 68.6 (#186), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkGemma 3 12BMistral Large
LMArena Instruction Following12991249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Gemma 3 12B leads

Gemma 3 12B: 40.0 (#162), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkGemma 3 12BMistral Large
LMArena Longer Query13171261

Writing & Preference Gemma 3 12B leads

Gemma 3 12B: 47.5 (#209), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkGemma 3 12BMistral Large
LMArena Text13341266
LMArena Creative Writing13311243
EQ-Bench Creative Writing1126985
LMArena Multi-Turn13341260
Short-Story Creative Writing—69%
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Gemma 3 12B better than Mistral Large?

Gemma 3 12B and Mistral Large score almost the same on the Noometry Index (32.1 vs 31.9), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 3 12B or Mistral Large?

Gemma 3 12B is cheaper. It lists at $0.05 per million input tokens and $0.15 per million output tokens; Mistral Large lists at $2 and $6.

Is Gemma 3 12B or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 31.7 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Gemma 3 12B and Mistral Large share?

22 benchmarks have published results for both models. Gemma 3 12B has 24 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper