Meta, open weights

Llama 2-13B

Llama 2-13B by Meta ranks 309th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.6. Its strongest category is math, where it ranks 229th.

Last verified

Specifications

Noometry rank
#309 of 354
Index score
29.6
Evidence
Confirmed 32 results
Provider
Meta
Released
July 18, 2023
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Llama 2-13B category scores
  1. Coding 30.9
  2. Reasoning 12.8
  3. Math 31.1
  4. Knowledge 28.1
  5. Multilingual 26.5
  6. Instruction Following 53.3
  7. Long Context 32.3
  8. Writing & Preference 29.8
Llama 2-13B category ranks
CategoryScoreRankResults
Coding30.9#2911
Reasoning12.8#3373
Math31.1#2291
Knowledge28.1#2491
Multilingual26.5#2791
Instruction Following53.3#2871
Long Context32.3#2691
Writing & Preference29.8#2893

Strengths and weaknesses

Categories where Llama 2-13B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 2-13B: strongest categories
CategoryScorevs medianRank
Math31.1−5.4#229 of 327, top 71%
Knowledge28.1−9.3#249 of 314, top 80%
Coding30.9−7.8#291 of 340, top 86%

Weakest categories

Llama 2-13B: weakest categories
CategoryScorevs medianRank
Reasoning12.8−10.8#337 of 350, top 97%
Instruction Following53.3−18.0#287 of 305, top 95%
Multilingual26.5−20.9#279 of 297, top 94%

Closest competitors

The models ranked just above and below Llama 2-13B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 2-13B
ModelRankScoreBlended $/MSpeed
Phi 3 Mini 128k Instruct#30529.7——Compare
phi-3-medium 14B#30629.7——Compare
Gemma 2B#30729.6——Compare
Llama 3.1-70B#30829.6$0.40—Compare
Claude 3 Opus#31029.5——Compare
DBRX#31129.4——Compare
Gemma 2 27B#31229.4$0.65—Compare
Gemma 1.1 2b IT#31329.3——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 2-13B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1062#281 of 294, top 96%LMArena2026-10-08

Reasoning

Llama 2-13B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Chess Puzzles0%#118 of 129, top 92%Epoch AI2026-08-28
LMArena Hard Prompts1051#281 of 297, top 95%LMArena2026-10-08
DTBench42.2%#150 of 151, top 100%Epoch AI
BIG-Bench Hard58.2%#15 of 27, top 56%Epoch AI
BIG-Bench Hard47%Epoch AI
Epoch Capabilities Index106.17#196 of 213, top 93%Epoch AI2023-07-18
HellaSwag80.7%#17 of 29, top 59%Epoch AI
LAMBADA76.5%#4 of 9, top 45%Epoch AI
PIQA80.8%#20 of 27, top 75%Epoch AI
WinoGrande72.8%#29 of 43, top 68%Epoch AI

Math

Llama 2-13B Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1065#275 of 285, top 97%LMArena2026-10-08
GSM8K34.3%Epoch AI
GSM8K36.9%#28 of 38, top 74%Epoch AI

Knowledge

Llama 2-13B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1030#268 of 273, top 99%LMArena2026-10-08
ARC (AI2) Challenge60.3%#23 of 39, top 59%Epoch AI
BoolQ82.4%#15 of 23, top 66%Epoch AI
MMLU50.9%Epoch AI
MMLU55.6%#68 of 81, top 84%Epoch AI
OpenBookQA57%#14 of 19, top 74%Epoch AI
TriviaQA79.6%#12 of 25, top 48%Epoch AI

Multimodal

Llama 2-13B Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
ScienceQA55.8%#4 of 6, top 67%Epoch AI

Multilingual

Llama 2-13B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1024#279 of 297, top 94%LMArena2026-10-08
LMArena Chinese1001#278 of 285, top 98%LMArena2026-10-08
LMArena French1044#219 of 223, top 99%LMArena2026-10-08
LMArena German1009#226 of 231, top 98%LMArena2026-10-08
LMArena Japanese894#210 of 211, top 100%LMArena2026-10-08
LMArena Korean953#208 of 213, top 98%LMArena2026-10-08
LMArena Russian1055#266 of 283, top 94%LMArena2026-10-08
LMArena Spanish1087#218 of 226, top 97%LMArena2026-10-08

Instruction Following

Llama 2-13B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1045#284 of 298, top 96%LMArena2026-10-08

Long Context

Llama 2-13B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1064#274 of 291, top 95%LMArena2026-10-08

Writing & Preference

Llama 2-13B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1084#277 of 297, top 94%LMArena2026-10-08
LMArena Creative Writing1047#278 of 295, top 95%LMArena2026-10-08
LMArena Multi-Turn1050#277 of 295, top 94%LMArena2026-10-08

Compare Llama 2-13B

Other Meta models

Frequently asked questions

How good is Llama 2-13B?

Llama 2-13B by Meta ranks 309th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.6. Its strongest category is math, where it ranks 229th.

Is Llama 2-13B open source?

Yes. Llama 2-13B's weights are downloadable; check the license for commercial terms.

What are Llama 2-13B's strengths and weaknesses?

Relative to other ranked models, Llama 2-13B places best in math, knowledge, coding and lowest in reasoning, instruction following, multilingual.

What is Llama 2-13B best at?

Its best category is math, where it ranks 229th on Noometry.