Google, open weights
Gemma 3 12B
Gemma 3 12B by Google ranks 262nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 32.1. Its strongest category is agentic & tool use, where it ranks 108th. API pricing starts at $0.05 per million input tokens and $0.15 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #262 of 354
- Index score
- 32.1
- Evidence
- Confirmed 24 results
- Provider
Google
- Released
- March 12, 2025
- Weights
- Open weights
- Reasoning
- No
- Context window
- 131K
- Max output
- 8K
- Input price
- $0.05 / M
- Output price
- $0.15 / M
- Blended price
- $0.075 / M
- Output speed
- Not measured
- Value
- #11 of 219
- Knowledge cutoff
- August 2024
- Input
- text, image
- Hugging Face
- google/gemma-3-12b-it
Category scores
Each category score combines every public result we have in that category.
- Coding 31.7
- Agentic & Tool Use 25.5
- Reasoning 15.7
- Math 22.3
- Knowledge 26.5
- Multilingual 45.7
- Instruction Following 68.6
- Long Context 40.0
- Writing & Preference 47.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 31.7 | #280 | 2 |
| Agentic & Tool Use | 25.5 | #108 | 1 |
| Reasoning | 15.7 | #313 | 5 |
| Math | 22.3 | #279 | 2 |
| Knowledge | 26.5 | #257 | 3 |
| Multilingual | 45.7 | #165 | 1 |
| Instruction Following | 68.6 | #186 | 1 |
| Long Context | 40.0 | #162 | 1 |
| Writing & Preference | 47.5 | #209 | 4 |
Strengths and weaknesses
Categories where Gemma 3 12B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 40.0 | −0.9 | #162 of 296, top 55% |
| Multilingual | 45.7 | −1.7 | #165 of 297, top 56% |
| Instruction Following | 68.6 | −2.7 | #186 of 305, top 61% |
Closest competitors
The models ranked just above and below Gemma 3 12B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Granite 3.1 8b Instruct | #258 | 32.4 | — | — | Compare |
| Pixtral Large | #259 | 32.2 | $3 | — | Compare |
| Falcon-180B | #260 | 32.2 | — | — | Compare |
| Gemini 1.5 Pro (May 2024) | #261 | 32.1 | — | — | Compare |
| Mistral Large | #263 | 31.9 | $3 | — | Compare |
| Qwen3-4B | #264 | 31.9 | — | — | Compare |
| Amazon Nova Lite | #265 | 31.9 | $0.10 | — | Compare |
| Mistral Medium 3.1 | #266 | 31.9 | $0.80 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SciCode | 17.4% | #118 of 121, top 98% | Epoch AI | ||
| LMArena Coding | 1281 | #210 of 294, top 72% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 30.4% | #33 of 49, top 68% | prompt | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CritPt | 0% | #107 of 134, top 80% | Epoch AI | ||
| Chess Puzzles | 0% | #110 of 129, top 86% | Epoch AI | 2026-08-28 | |
| LMArena Hard Prompts | 1309 | #183 of 297, top 62% | LMArena | 2026-10-08 | |
| DTBench | 48.8% | #139 of 151, top 93% | Epoch AI | ||
| LMCA | 4.5% | #124 of 125, top 100% | Epoch AI | ||
| Epoch Capabilities Index | 123.5 | #160 of 213, top 76% | Epoch AI | 2025-03-12 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 16.7% | #128 of 173, top 74% | Epoch AI | 2026-08-28 | |
| LMArena Math | 1307 | #178 of 285, top 63% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 39.5% | #151 of 186, top 82% | Epoch AI | 2026-08-28 | |
| Vectara Hallucination Rate (lower is better) | 4.4% | #5 of 96, top 6% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1248 | #198 of 273, top 73% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| MindCube | 46.7% | Best of 2 | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1318 | #165 of 297, top 56% | LMArena | 2026-10-08 | |
| LMArena German | 1370 | #116 of 231, top 51% | LMArena | 2026-10-08 | |
| LMArena Russian | 1335 | #155 of 283, top 55% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1299 | #179 of 298, top 61% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1317 | #174 of 291, top 60% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1334 | #170 of 297, top 58% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1331 | #145 of 295, top 50% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1126 | #92 of 115, top 80% | EQ-Bench | ||
| LMArena Multi-Turn | 1334 | #169 of 295, top 58% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| bedrock | $0.09 | $0.29 | — | 2026-10-10 |
| deepinfra | $0.05 | $0.15 | — | 2026-10-10 |
| openrouter | $0.05 | $0.15 | — | 2026-10-10 |
Compare Gemma 3 12B
- Gemma 3 12B vs Gemma 3 27B
- Gemma 3 12B vs Gemini 1.5 Pro (May 2024)
- Gemma 3 12B vs Mistral Large
- Gemma 3 12B vs Falcon-180B
- Gemma 3 12B vs Qwen3-4B
- Gemma 3 12B vs Pixtral Large
- Gemma 3 12B vs Amazon Nova Lite
- Gemma 3 12B vs GPT-6 Astra
- Gemma 3 12B vs Claude Fable 5.1
- Gemma 3 12B vs Kimi K3
- Gemma 3 12B vs Grok 4.6
- Gemma 3 12B vs Qwen3.8 Max
- Gemma 3 12B vs GLM-5.3
- Gemma 3 12B vs Muse Spark 1.3
Other Google models
Frequently asked questions
How good is Gemma 3 12B?
Gemma 3 12B by Google ranks 262nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 32.1. Its strongest category is agentic & tool use, where it ranks 108th. API pricing starts at $0.05 per million input tokens and $0.15 per million output tokens, with a 131K-token context window.
How much does Gemma 3 12B cost?
Gemma 3 12B costs $0.05 per million input tokens and $0.15 per million output tokens on deepinfra.
What is Gemma 3 12B's context window?
Gemma 3 12B accepts up to 131K tokens of input and can write up to 8K tokens in one response.
Is Gemma 3 12B open source?
Yes. Gemma 3 12B's weights are downloadable from Hugging Face (google/gemma-3-12b-it); check the license for commercial terms.
What are Gemma 3 12B's strengths and weaknesses?
Relative to other ranked models, Gemma 3 12B places best in long context, multilingual, instruction following and lowest in reasoning, math, coding.
What is Gemma 3 12B best at?
Its best category is agentic & tool use, where it ranks 108th on Noometry.