Meta, open weights

Llama 2-70B

Llama 2-70B by Meta ranks 349th of 354 ranked models on the Noometry Index as of October 2026, with a score of 24.4. Its strongest category is long context, where it ranks 270th.

Last verified

Specifications

Noometry rank
#349 of 354
Index score
24.4
Evidence
Confirmed 35 results
Provider
Meta
Released
July 18, 2023
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Llama 2-70B category scores
  1. Coding 31.4
  2. Reasoning 14.4
  3. Math 8.1
  4. Knowledge 7.4
  5. Multilingual 27.7
  6. Instruction Following 54.9
  7. Long Context 32.3
  8. Writing & Preference 32.3
Llama 2-70B category ranks
CategoryScoreRankResults
Coding31.4#2861
Reasoning14.4#3252
Math8.1#3263
Knowledge7.4#3102
Multilingual27.7#2741
Instruction Following54.9#2781
Long Context32.3#2701
Writing & Preference32.3#2793

Strengths and weaknesses

Categories where Llama 2-70B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 2-70B: strongest categories
CategoryScorevs medianRank
Coding31.4−7.3#286 of 340, top 85%
Writing & Preference32.3−21.5#279 of 312, top 90%
Instruction Following54.9−16.4#278 of 305, top 92%

Weakest categories

Llama 2-70B: weakest categories
CategoryScorevs medianRank
Math8.1−28.5#326 of 327, top 100%
Knowledge7.4−29.9#310 of 314, top 99%
Reasoning14.4−9.2#325 of 350, top 93%

Closest competitors

The models ranked just above and below Llama 2-70B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 2-70B
ModelRankScoreBlended $/MSpeed
Claude 2.1#34525.2——Compare
Claude 2#34625.0——Compare
DeepSeek LLM 67B#34724.9——Compare
Llama 13b#34824.4——Compare
GPT-3.5-turbo#35023.2$0.75—Compare
Mistral 7B#35123.0$0.25—Compare
Llama 3.1-8B#35223.0$0.0575—Compare
Gemma 3 1B#35321.1——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 2-70B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1079#278 of 294, top 95%LMArena2026-10-08

Reasoning

Llama 2-70B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1073#274 of 297, top 93%LMArena2026-10-08
DTBench41.6%#151 of 151, top 100%Epoch AI
BIG-Bench Hard58.5%Epoch AI
BIG-Bench Hard64.9%#11 of 27, top 41%Epoch AI
CommonsenseQA 2.050%#2 of 2Epoch AI
Epoch Capabilities Index113.79#185 of 213, top 87%Epoch AI2023-07-18
ForecastBench51.4#71 of 72, top 99%Epoch AI
HellaSwag85.3%#8 of 29, top 28%Epoch AI
LAMBADA78.9%#2 of 9, top 23%Epoch AI
PIQA82.8%#13 of 27, top 49%Epoch AI
WinoGrande80.2%#15 of 43, top 35%Epoch AI

Math

Llama 2-70B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20250%#173 of 173, top 100%Epoch AI2025-02-25
LMArena Math1091#267 of 285, top 94%LMArena2026-10-08
MATH Level 53.3%#79 of 79, top 100%Epoch AI2025-01-27
GSM8K69.6%#15 of 38, top 40%Epoch AI
GSM8K58.7%Epoch AI

Knowledge

Llama 2-70B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond26.3%#178 of 186, top 96%Epoch AI2025-01-27
LMArena Expert1039#265 of 273, top 98%LMArena2026-10-08
ARC (AI2) Challenge78.3%#15 of 39, top 39%Epoch AI
BoolQ88.6%#3 of 23, top 14%Epoch AI
MMLU69.9%#46 of 81, top 57%Epoch AI
MMLU59.9%Epoch AI
OpenBookQA60.2%#11 of 19, top 58%Epoch AI
TriviaQA87.6%Best of 25Epoch AI

Multilingual

Llama 2-70B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1045#274 of 297, top 93%LMArena2026-10-08
LMArena Chinese995#279 of 285, top 98%LMArena2026-10-08
LMArena French1090#215 of 223, top 97%LMArena2026-10-08
LMArena German1041#223 of 231, top 97%LMArena2026-10-08
LMArena Japanese927#208 of 211, top 99%LMArena2026-10-08
LMArena Korean964#205 of 213, top 97%LMArena2026-10-08
LMArena Russian1083#261 of 283, top 93%LMArena2026-10-08
LMArena Spanish1143#207 of 226, top 92%LMArena2026-10-08

Instruction Following

Llama 2-70B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1071#274 of 298, top 92%LMArena2026-10-08

Long Context

Llama 2-70B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1062#275 of 291, top 95%LMArena2026-10-08

Writing & Preference

Llama 2-70B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1115#269 of 297, top 91%LMArena2026-10-08
LMArena Creative Writing1075#273 of 295, top 93%LMArena2026-10-08
LMArena Multi-Turn1088#268 of 295, top 91%LMArena2026-10-08

Compare Llama 2-70B

Other Meta models

Frequently asked questions

How good is Llama 2-70B?

Llama 2-70B by Meta ranks 349th of 354 ranked models on the Noometry Index as of October 2026, with a score of 24.4. Its strongest category is long context, where it ranks 270th.

Is Llama 2-70B open source?

Yes. Llama 2-70B's weights are downloadable; check the license for commercial terms.

What are Llama 2-70B's strengths and weaknesses?

Relative to other ranked models, Llama 2-70B places best in coding, writing & preference, instruction following and lowest in math, knowledge, reasoning.

What is Llama 2-70B best at?

Its best category is long context, where it ranks 270th on Noometry.