Model comparison

Gemma 3n E4b IT vs o3-mini

Gemma 3n E4b IT and o3-mini score almost the same on the Noometry Index (37.3 vs 36.7), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Gemma 3n E4b IT Google

37.3

Rank #206 Confirmed

o3-mini OpenAI

36.7

Rank #212 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 3n E4b IT scores higher in 3 categories and o3-mini in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where o3-mini leads 75.1 to 66.1.
  • Gemma 3n E4b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 3n E4b IT and o3-mini specifications
Gemma 3n E4b ITo3-mini
ProviderGoogleOpenAI
Noometry Index37.336.7
Released—2024-12-20
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$1.10
Output $ / M tokens—$4.40
Results tracked1851

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-mini leads

Gemma 3n E4b IT: 37.0 (#198), o3-mini: 40.8 (#132)

Coding benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Coding12681378
Aider Polyglot—60.4%
SciCode—39.8%
GSO—1.3%
WeirdML—43.7%
LiveBench Coding—82.7%
CadEval—54%

Agentic & Tool Use Not comparable

Gemma 3n E4b IT: —, o3-mini: 29.6 (#84)

Agentic & Tool Use benchmarks
BenchmarkGemma 3n E4b ITo3-mini
Cybench—22.5%

Reasoning Gemma 3n E4b IT leads

Gemma 3n E4b IT: 19.9 (#247), o3-mini: 16.3 (#305)

Reasoning benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Hard Prompts12841366
ARC-AGI-2—3%
SimpleBench—22.8%
Kagi LLM Benchmark31.5%—
ARC-AGI-1—34.5%
CritPt—0.3%
Chess Puzzles—17%
LiveBench Reasoning—89.6%
Mystery Game Puzzles—7%
DTBench—68.8%
LiveBench Data Analysis—70.6%
LMCA—19%
Epoch Capabilities Index—140.34
ForecastBench—59.6
LiveBench—75.9%

Math Gemma 3n E4b IT leads

Gemma 3n E4b IT: 35.1 (#188), o3-mini: 28.1 (#244)

Math benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Math12511396
FrontierMath (Tiers 1-3)—18.6%
FrontierMath Tier 4—0%
OTIS Mock AIME 2024-2025—76.9%
LiveBench Math—77.3%
MATH Level 5—96.5%
FrontierMath (Feb 2025 set)—12.4%
FrontierMath Tier 4 (v1)—4.2%

Knowledge o3-mini leads

Gemma 3n E4b IT: 34.2 (#198), o3-mini: 38.3 (#146)

Knowledge benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Expert12461364
GPQA Diamond—77%
SimpleQA Verified—15.3%
Confabulations—17.9%

Multilingual o3-mini leads

Gemma 3n E4b IT: 43.4 (#183), o3-mini: 45.7 (#164)

Multilingual benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Non-English12851319
LMArena Chinese13091379
LMArena French13301334
LMArena German13111303
LMArena Japanese12721286
LMArena Korean12591314
LMArena Russian12881304
LMArena Spanish13051321

Instruction Following o3-mini leads

Gemma 3n E4b IT: 66.1 (#210), o3-mini: 75.1 (#72)

Instruction Following benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Instruction Following12551337
LiveBench Instruction Following—84.4%

Long Context Gemma 3n E4b IT leads

Gemma 3n E4b IT: 38.7 (#191), o3-mini: 33.8 (#256)

Long Context benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Longer Query12761343
Fiction.LiveBench—50%

Writing & Preference Too close to call

Gemma 3n E4b IT: 50.1 (#186), o3-mini: 50.3 (#182)

Writing & Preference benchmarks
BenchmarkGemma 3n E4b ITo3-mini
LMArena Text13061337
LMArena Creative Writing12871286
LMArena Multi-Turn12761320
Short-Story Creative Writing—61.7%
LiveBench Language—50.7%

Frequently asked questions

Is Gemma 3n E4b IT better than o3-mini?

Gemma 3n E4b IT and o3-mini score almost the same on the Noometry Index (37.3 vs 36.7), so choose on price, context window or the category you care about most.

Is Gemma 3n E4b IT or o3-mini better for coding?

o3-mini scores higher on coding benchmarks: 40.8 versus 37.0 in the Noometry coding category.

How many benchmarks do Gemma 3n E4b IT and o3-mini share?

17 benchmarks have published results for both models. Gemma 3n E4b IT has 18 scored results on Noometry and o3-mini has 51.

Related comparisons

Go deeper