Meta, open weights
Llama 2-7B
Llama 2-7B by Meta ranks 317th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.1. Its strongest category is math, where it ranks 233rd.
Last verified
Specifications
- Noometry rank
- #317 of 354
- Index score
- 29.1
- Evidence
- Confirmed 29 results
- Provider
Meta
- Released
- July 18, 2023
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 29.2
- Reasoning 15.7
- Math 30.7
- Knowledge 28.2
- Multilingual 23.8
- Instruction Following 50.8
- Long Context 30.4
- Writing & Preference 28.0
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 29.2 | #307 | 1 |
| Reasoning | 15.7 | #312 | 2 |
| Math | 30.7 | #233 | 1 |
| Knowledge | 28.2 | #248 | 1 |
| Multilingual | 23.8 | #293 | 1 |
| Instruction Following | 50.8 | #298 | 1 |
| Long Context | 30.4 | #287 | 1 |
| Writing & Preference | 28.0 | #298 | 3 |
Strengths and weaknesses
Categories where Llama 2-7B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 23.8 | −23.7 | #293 of 297, top 99% |
| Instruction Following | 50.8 | −20.5 | #298 of 305, top 98% |
| Long Context | 30.4 | −10.5 | #287 of 296, top 97% |
Closest competitors
The models ranked just above and below Llama 2-7B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Gemma 1.1 2b IT | #313 | 29.3 | — | — | Compare |
| Phi 3 Small 8k Instruct | #314 | 29.3 | — | — | Compare |
| Claude 3.5 Haiku | #315 | 29.2 | — | — | Compare |
| GPT-4 | #316 | 29.1 | $37.50 | — | Compare |
| Granite 4.0 Micro | #318 | 29.0 | $0.0408 | — | Compare |
| Claude 3 Sonnet | #319 | 29.0 | — | — | Compare |
| Qwen2.5 7B Instruct | #320 | 29.0 | $0.31 | — | Compare |
| Llama 3.2 3B | #321 | 28.9 | $0.12 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 1002 | #290 of 294, top 99% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Chess Puzzles | 0% | #119 of 129, top 93% | Epoch AI | 2026-08-30 | |
| LMArena Hard Prompts | 1009 | #289 of 297, top 98% | LMArena | 2026-10-08 | |
| BIG-Bench Hard | 39.2% | #22 of 27, top 82% | Epoch AI | ||
| Epoch Capabilities Index | 99.06 | #204 of 213, top 96% | Epoch AI | 2023-07-18 | |
| HellaSwag | 77.2% | #20 of 29, top 69% | Epoch AI | ||
| LAMBADA | 73.3% | #7 of 9, top 78% | Epoch AI | ||
| PIQA | 78.8% | #24 of 27, top 89% | Epoch AI | ||
| WinoGrande | 69.2% | #32 of 43, top 75% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 1042 | #279 of 285, top 98% | LMArena | 2026-10-08 | |
| GSM8K | 16.7% | #36 of 38, top 95% | Epoch AI |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Expert | 1036 | #266 of 273, top 98% | LMArena | 2026-10-08 | |
| ARC (AI2) Challenge | 45.9% | #31 of 39, top 80% | Epoch AI | ||
| BoolQ | 77.9% | #18 of 23, top 79% | Epoch AI | ||
| MMLU | 45.8% | #71 of 81, top 88% | Epoch AI | ||
| OpenBookQA | 58.6% | #12 of 19, top 64% | Epoch AI | ||
| TriviaQA | 73.7% | #17 of 25, top 68% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ScienceQA | 43.1% | #6 of 6, top 100% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 973 | #293 of 297, top 99% | LMArena | 2026-10-08 | |
| LMArena Chinese | 973 | #282 of 285, top 99% | LMArena | 2026-10-08 | |
| LMArena French | 970 | #223 of 223, top 100% | LMArena | 2026-10-08 | |
| LMArena German | 978 | #229 of 231, top 100% | LMArena | 2026-10-08 | |
| LMArena Russian | 995 | #276 of 283, top 98% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1007 | #226 of 226, top 100% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1006 | #292 of 298, top 98% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 999 | #287 of 291, top 99% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1053 | #288 of 297, top 97% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1033 | #283 of 295, top 96% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1029 | #282 of 295, top 96% | LMArena | 2026-10-08 |
Compare Llama 2-7B
- Llama 2-7B vs Llama 13b
- Llama 2-7B vs GPT-4
- Llama 2-7B vs Granite 4.0 Micro
- Llama 2-7B vs Claude 3.5 Haiku
- Llama 2-7B vs Claude 3 Sonnet
- Llama 2-7B vs Phi 3 Small 8k Instruct
- Llama 2-7B vs Qwen2.5 7B Instruct
- Llama 2-7B vs GPT-6 Astra
- Llama 2-7B vs Claude Fable 5.1
- Llama 2-7B vs Gemini 3.8 Flash
- Llama 2-7B vs Kimi K3
- Llama 2-7B vs Grok 4.6
- Llama 2-7B vs Qwen3.8 Max
- Llama 2-7B vs GLM-5.3
Other Meta models
- Muse Spark 1.354.8
- Muse Spark50.6
- Muse Spark 1.250.3
- Muse Spark 1.149.9
- Muse Glimmer41.7
- Codellama 70b Instruct33.7
- Llama 4 Maverick30.9
- Codellama 34b Instruct30.8
Frequently asked questions
How good is Llama 2-7B?
Llama 2-7B by Meta ranks 317th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.1. Its strongest category is math, where it ranks 233rd.
Is Llama 2-7B open source?
Yes. Llama 2-7B's weights are downloadable; check the license for commercial terms.
What are Llama 2-7B's strengths and weaknesses?
Relative to other ranked models, Llama 2-7B places best in math, knowledge, reasoning and lowest in multilingual, instruction following, long context.
What is Llama 2-7B best at?
Its best category is math, where it ranks 233rd on Noometry.