Meta, open weights
Llama 3.1-8B
Llama 3.1-8B by Meta ranks 352nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.0. Its strongest category is agentic & tool use, where it ranks 131st. API pricing starts at $0.05 per million input tokens and $0.08 per million output tokens, with a 128K-token context window.
Last verified
Specifications
- Noometry rank
- #352 of 354
- Index score
- 23.0
- Evidence
- Confirmed 43 results
- Provider
Meta
- Released
- July 23, 2024
- Weights
- Open weights
- Reasoning
- No
- Context window
- 128K
- Max output
- 4K
- Input price
- $0.05 / M
- Output price
- $0.08 / M
- Blended price
- $0.0575 / M
- Output speed
- Not measured
- Value
- #13 of 219
- Knowledge cutoff
- December 2023
- Input
- text
- Hugging Face
- meta-llama/Meta-Llama-3.1-8B-Instruct
Category scores
Each category score combines every public result we have in that category.
- Coding 20.2
- Agentic & Tool Use 22.5
- Reasoning 14.9
- Math 10.2
- Knowledge 8.0
- Multilingual 34.0
- Instruction Following 58.9
- Long Context 35.8
- Writing & Preference 29.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 20.2 | #340 | 5 |
| Agentic & Tool Use | 22.5 | #131 | 2 |
| Reasoning | 14.9 | #321 | 5 |
| Math | 10.2 | #317 | 4 |
| Knowledge | 8.0 | #307 | 4 |
| Multilingual | 34.0 | #249 | 1 |
| Instruction Following | 58.9 | #258 | 2 |
| Long Context | 35.8 | #238 | 1 |
| Writing & Preference | 29.7 | #290 | 5 |
Strengths and weaknesses
Categories where Llama 3.1-8B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 35.8 | −5.1 | #238 of 296, top 81% |
| Multilingual | 34.0 | −13.4 | #249 of 297, top 84% |
| Instruction Following | 58.9 | −12.4 | #258 of 305, top 85% |
Closest competitors
The models ranked just above and below Llama 3.1-8B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude 2 | #346 | 25.0 | — | — | Compare |
| DeepSeek LLM 67B | #347 | 24.9 | — | — | Compare |
| Llama 13b | #348 | 24.4 | — | — | Compare |
| Llama 2-70B | #349 | 24.4 | — | — | Compare |
| GPT-3.5-turbo | #350 | 23.2 | $0.75 | — | Compare |
| Mistral 7B | #351 | 23.0 | $0.25 | — | Compare |
| Gemma 3 1B | #353 | 21.1 | — | — | Compare |
| Llama 3.2 1B | #354 | 20.1 | $0.0705 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SciCode | 13.2% | #120 of 121, top 100% | Epoch AI | ||
| WeirdML | 1.7% | #119 of 119, top 100% | Epoch AI | ||
| BigCodeBench Instruct | 32.8% | #51 of 64, top 80% | BigCodeBench | 2024-07-23 | |
| LMArena Coding | 1195 | #244 of 294, top 83% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 40.5% | #52 of 66, top 79% | BigCodeBench | 2024-07-23 | |
| HumanEval+ | 62.8% | #25 of 45, top 56% | EvalPlus | ||
| MBPP+ | 55.6% | #28 of 38, top 74% | EvalPlus |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 25.8% | #41 of 49, top 84% | prompt | Berkeley Function Calling Leaderboard | |
| BALROG | 15.1% | #30 of 35, top 86% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CritPt | 0% | #116 of 134, top 87% | Epoch AI | ||
| Chess Puzzles | 0% | #120 of 129, top 94% | Epoch AI | 2026-08-27 | |
| LMArena Hard Prompts | 1175 | #245 of 297, top 83% | LMArena | 2026-10-08 | |
| DTBench | 50.9% | #135 of 151, top 90% | Epoch AI | ||
| LMCA | 5.4% | #122 of 125, top 98% | Epoch AI | ||
| Epoch Capabilities Index | 116.57 | #182 of 213, top 86% | Epoch AI | 2024-07-23 | |
| PIQA | 81.2% | #18 of 27, top 67% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 1.7% | #164 of 173, top 95% | Epoch AI | 2026-08-27 | |
| Omni-MATH | 13.7% | #55 of 57, top 97% | HELM Capabilities | ||
| LMArena Math | 1179 | #242 of 285, top 85% | LMArena | 2026-10-08 | |
| MATH Level 5 | 22.9% | #61 of 79, top 78% | Epoch AI | 2025-01-27 | |
| GSM8K | 82.4% | #12 of 38, top 32% | Epoch AI |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 27% | #177 of 186, top 96% | Epoch AI | 2026-08-27 | |
| MMLU-Pro | 40.6% | #54 of 58, top 94% | HELM Capabilities | ||
| GPQA (HELM) | 24.7% | #57 of 57, top 100% | HELM Capabilities | ||
| LMArena Expert | 1144 | #239 of 273, top 88% | LMArena | 2026-10-08 | |
| BoolQ | 82.8% | #13 of 23, top 57% | Epoch AI | ||
| MMLU | 56.1% | #67 of 81, top 83% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1148 | #249 of 297, top 84% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1151 | #246 of 285, top 87% | LMArena | 2026-10-08 | |
| LMArena French | 1177 | #198 of 223, top 89% | LMArena | 2026-10-08 | |
| LMArena German | 1144 | #203 of 231, top 88% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1061 | #190 of 211, top 91% | LMArena | 2026-10-08 | |
| LMArena Korean | 1053 | #193 of 213, top 91% | LMArena | 2026-10-08 | |
| LMArena Russian | 1158 | #248 of 283, top 88% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1169 | #200 of 226, top 89% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 74.3% | #51 of 57, top 90% | HELM Capabilities | ||
| LMArena Instruction Following | 1159 | #250 of 298, top 84% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1182 | #247 of 291, top 85% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1187 | #249 of 297, top 84% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1154 | #249 of 295, top 85% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 713 | #109 of 115, top 95% | EQ-Bench | ||
| WildBench | 68.7% | #53 of 57, top 93% | HELM Capabilities | ||
| LMArena Multi-Turn | 1172 | #245 of 295, top 84% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| bedrock | $0.22 | $0.22 | — | 2026-10-10 |
| groq | $0.05 | $0.08 | — | 2026-10-10 |
| openrouter | $0.05 | $0.08 | $0.025 | 2026-10-10 |
Compare Llama 3.1-8B
- Llama 3.1-8B vs Llama 3-70B
- Llama 3.1-8B vs Mistral 7B
- Llama 3.1-8B vs Gemma 3 1B
- Llama 3.1-8B vs GPT-3.5-turbo
- Llama 3.1-8B vs Llama 3.2 1B
- Llama 3.1-8B vs Llama 2-70B
- Llama 3.1-8B vs Llama 13b
- Llama 3.1-8B vs GPT-6 Astra
- Llama 3.1-8B vs Claude Fable 5.1
- Llama 3.1-8B vs Gemini 3.8 Flash
- Llama 3.1-8B vs Kimi K3
- Llama 3.1-8B vs Grok 4.6
- Llama 3.1-8B vs Qwen3.8 Max
- Llama 3.1-8B vs GLM-5.3
Other Meta models
- Muse Spark 1.354.8
- Muse Spark50.6
- Muse Spark 1.250.3
- Muse Spark 1.149.9
- Muse Glimmer41.7
- Codellama 70b Instruct33.7
- Llama 4 Maverick30.9
- Codellama 34b Instruct30.8
Frequently asked questions
How good is Llama 3.1-8B?
Llama 3.1-8B by Meta ranks 352nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.0. Its strongest category is agentic & tool use, where it ranks 131st. API pricing starts at $0.05 per million input tokens and $0.08 per million output tokens, with a 128K-token context window.
How much does Llama 3.1-8B cost?
Llama 3.1-8B costs $0.05 per million input tokens and $0.08 per million output tokens on groq.
What is Llama 3.1-8B's context window?
Llama 3.1-8B accepts up to 128K tokens of input and can write up to 4K tokens in one response.
Is Llama 3.1-8B open source?
Yes. Llama 3.1-8B's weights are downloadable from Hugging Face (meta-llama/Meta-Llama-3.1-8B-Instruct); check the license for commercial terms.
What are Llama 3.1-8B's strengths and weaknesses?
Relative to other ranked models, Llama 3.1-8B places best in long context, multilingual, instruction following and lowest in coding, knowledge, math.
What is Llama 3.1-8B best at?
Its best category is agentic & tool use, where it ranks 131st on Noometry.