Model comparison

Codestral vs Gemma 3 4B

Codestral is the stronger model overall, scoring 30.6 to 28.1 on the Noometry Index. Gemma 3 4B costs 9.0× less per token, which makes it the better buy when Codestral's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Gemma 3 4B Google

28.1

Rank #326 Confirmed

Summary

  • They share 1 benchmark with published results for both. Codestral scores higher in 1 category and Gemma 3 4B in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Gemma 3 4B leads 35.9 to 27.3.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 32.5% for Codestral and 25.2% for Gemma 3 4B.
  • Gemma 3 4B is cheaper at $0.04 / $0.08 per million input/output tokens, against $0.30 / $0.90 for Codestral.
  • Codestral accepts more context: 256K tokens versus 131K.
  • Gemma 3 4B has downloadable open weights; the other is API-only.

Side by side

Codestral and Gemma 3 4B specifications
CodestralGemma 3 4B
ProviderMistral AIGoogle
Noometry Index30.628.1
Released2024-05-292025-03-12
WeightsProprietaryOpen
Context window256K131K
Max output8K4K
Input $ / M tokens$0.30$0.04
Output $ / M tokens$0.90$0.08
Results tracked722

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3 4B leads

Codestral: 27.3 (#321), Gemma 3 4B: 35.9 (#215)

Coding benchmarks
BenchmarkCodestralGemma 3 4B
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—1230
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, Gemma 3 4B: 20.9 (#142)

Agentic & Tool Use benchmarks
BenchmarkCodestralGemma 3 4B
Berkeley Function Calling Leaderboard—19.6%

Reasoning Codestral leads

Codestral: 19.8 (#251), Gemma 3 4B: 13.2 (#335)

Reasoning benchmarks
BenchmarkCodestralGemma 3 4B
Kagi LLM Benchmark32.5%25.2%
Chess Puzzles—0%
LMArena Hard Prompts—1253
DTBench—50.9%
LMCA—2.8%
Epoch Capabilities Index—116.02

Math Not comparable

Codestral: —, Gemma 3 4B: 16.8 (#292)

Math benchmarks
BenchmarkCodestralGemma 3 4B
OTIS Mock AIME 2024-2025—7.5%
LMArena Math—1239

Knowledge Not comparable

Codestral: —, Gemma 3 4B: 11.8 (#299)

Knowledge benchmarks
BenchmarkCodestralGemma 3 4B
GPQA Diamond—23.2%
Vectara Hallucination Rate—6.4%
LMArena Expert—1223

Multilingual Not comparable

Codestral: —, Gemma 3 4B: 42.5 (#194)

Multilingual benchmarks
BenchmarkCodestralGemma 3 4B
LMArena Non-English—1273
LMArena German—1281
LMArena Russian—1294

Instruction Following Not comparable

Codestral: —, Gemma 3 4B: 65.2 (#225)

Instruction Following benchmarks
BenchmarkCodestralGemma 3 4B
LMArena Instruction Following—1239

Long Context Not comparable

Codestral: —, Gemma 3 4B: 38.7 (#194)

Long Context benchmarks
BenchmarkCodestralGemma 3 4B
LMArena Longer Query—1273

Writing & Preference Not comparable

Codestral: —, Gemma 3 4B: 42.0 (#239)

Writing & Preference benchmarks
BenchmarkCodestralGemma 3 4B
LMArena Text—1291
LMArena Creative Writing—1271
EQ-Bench Creative Writing—1068
LMArena Multi-Turn—1255

Frequently asked questions

Is Codestral better than Gemma 3 4B?

Codestral is the stronger model overall, scoring 30.6 to 28.1 on the Noometry Index. Gemma 3 4B costs 9.0× less per token, which makes it the better buy when Codestral's lead doesn't matter for your workload.

Which is cheaper, Codestral or Gemma 3 4B?

Gemma 3 4B is cheaper. It lists at $0.04 per million input tokens and $0.08 per million output tokens; Codestral lists at $0.30 and $0.90.

Is Codestral or Gemma 3 4B better for coding?

Gemma 3 4B scores higher on coding benchmarks: 35.9 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

Codestral does, with 256K tokens against 131K.

How many benchmarks do Codestral and Gemma 3 4B share?

1 benchmark has published results for both models. Codestral has 7 scored results on Noometry and Gemma 3 4B has 22.

Related comparisons

Go deeper