Meta, open weights

Llama 3.1-70B

Llama 3.1-70B by Meta ranks 308th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.6. Its strongest category is agentic & tool use, where it ranks 112th. API pricing starts at $0.40 per million input tokens and $0.40 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#308 of 354
Index score
29.6
Evidence
Confirmed 35 results
Provider
Meta
Released
July 23, 2024
Weights
Open weights
Reasoning
No
Context window
128K
Max output
4K
Input price
$0.40 / M
Output price
$0.40 / M
Blended price
$0.40 / M
Output speed
Not measured
Value
#79 of 219
Knowledge cutoff
December 2023
Input
text

Category scores

Each category score combines every public result we have in that category.

Llama 3.1-70B category scores
  1. Coding 30.3
  2. Agentic & Tool Use 25.1
  3. Reasoning 21.6
  4. Math 13.5
  5. Knowledge 24.2
  6. Multilingual 38.8
  7. Instruction Following 65.3
  8. Long Context 37.6
  9. Writing & Preference 35.4
Llama 3.1-70B category ranks
CategoryScoreRankResults
Coding30.3#2964
Agentic & Tool Use25.1#1122
Reasoning21.6#2203
Math13.5#3044
Knowledge24.2#2694
Multilingual38.8#2251
Instruction Following65.3#2232
Long Context37.6#2141
Writing & Preference35.4#2675

Strengths and weaknesses

Categories where Llama 3.1-70B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 3.1-70B: strongest categories
CategoryScorevs medianRank
Reasoning21.6−2.0#220 of 350, top 63%
Long Context37.6−3.3#214 of 296, top 73%
Agentic & Tool Use25.1−5.3#112 of 154, top 73%

Weakest categories

Llama 3.1-70B: weakest categories
CategoryScorevs medianRank
Math13.5−23.1#304 of 327, top 93%
Coding30.3−8.5#296 of 340, top 88%
Knowledge24.2−13.1#269 of 314, top 86%

Closest competitors

The models ranked just above and below Llama 3.1-70B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 3.1-70B
ModelRankScoreBlended $/MSpeed
OLMo 2 Furious 13B#30429.7——Compare
Phi 3 Mini 128k Instruct#30529.7——Compare
phi-3-medium 14B#30629.7——Compare
Gemma 2B#30729.6——Compare
Llama 2-13B#30929.6——Compare
Claude 3 Opus#31029.5——Compare
DBRX#31129.4——Compare
Gemma 2 27B#31229.4$0.65—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 3.1-70B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML9%#115 of 119, top 97%Epoch AI
BigCodeBench Instruct46.1%#14 of 64, top 22%BigCodeBench2024-07-23
LMArena Coding1260#222 of 294, top 76%LMArena2026-10-08
BigCodeBench Complete54.8%#20 of 66, top 31%BigCodeBench2024-07-23

Agentic & Tool Use

Llama 3.1-70B Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
TheAgentCompany6.9%#10 of 14, top 72%Epoch AI
BALROG27.9%#20 of 35, top 58%Epoch AI

Reasoning

Llama 3.1-70B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1241#226 of 297, top 77%LMArena2026-10-08
DTBench60%#116 of 151, top 77%Epoch AI
LMCA14.8%#105 of 125, top 84%Epoch AI
Epoch Capabilities Index125.92#154 of 213, top 73%Epoch AI2024-07-23

Math

Llama 3.1-70B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20253.6%#154 of 173, top 90%Epoch AI2025-02-25
Omni-MATH21%#50 of 57, top 88%HELM Capabilities
LMArena Math1252#210 of 285, top 74%LMArena2026-10-08
MATH Level 536.7%#55 of 79, top 70%Epoch AI2025-01-27

Knowledge

Llama 3.1-70B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond44.2%#143 of 186, top 77%Epoch AI2025-01-27
MMLU-Pro65.3%#38 of 58, top 66%HELM Capabilities
GPQA (HELM)42.6%#40 of 57, top 71%HELM Capabilities
LMArena Expert1209#218 of 273, top 80%LMArena2026-10-08
MMLU80.1%#16 of 81, top 20%Epoch AI

Multilingual

Llama 3.1-70B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1219#225 of 297, top 76%LMArena2026-10-08
LMArena Chinese1215#225 of 285, top 79%LMArena2026-10-08
LMArena French1261#179 of 223, top 81%LMArena2026-10-08
LMArena German1222#186 of 231, top 81%LMArena2026-10-08
LMArena Japanese1132#179 of 211, top 85%LMArena2026-10-08
LMArena Korean1140#182 of 213, top 86%LMArena2026-10-08
LMArena Russian1234#221 of 283, top 79%LMArena2026-10-08
LMArena Spanish1253#182 of 226, top 81%LMArena2026-10-08

Instruction Following

Llama 3.1-70B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval82.1%#33 of 57, top 58%HELM Capabilities
LMArena Instruction Following1231#227 of 298, top 77%LMArena2026-10-08

Long Context

Llama 3.1-70B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1241#226 of 291, top 78%LMArena2026-10-08

Writing & Preference

Llama 3.1-70B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1261#222 of 297, top 75%LMArena2026-10-08
LMArena Creative Writing1232#220 of 295, top 75%LMArena2026-10-08
EQ-Bench Creative Writing784#104 of 115, top 91%EQ-Bench
WildBench75.8%#44 of 57, top 78%HELM Capabilities
LMArena Multi-Turn1256#219 of 295, top 75%LMArena2026-10-08

API pricing by provider

Llama 3.1-70B API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$0.72$0.72—2026-10-10
openrouter$0.40$0.40—2026-10-10

Compare Llama 3.1-70B

Other Meta models

Frequently asked questions

How good is Llama 3.1-70B?

Llama 3.1-70B by Meta ranks 308th of 354 ranked models on the Noometry Index as of October 2026, with a score of 29.6. Its strongest category is agentic & tool use, where it ranks 112th. API pricing starts at $0.40 per million input tokens and $0.40 per million output tokens, with a 128K-token context window.

How much does Llama 3.1-70B cost?

Llama 3.1-70B costs $0.40 per million input tokens and $0.40 per million output tokens on openrouter.

What is Llama 3.1-70B's context window?

Llama 3.1-70B accepts up to 128K tokens of input and can write up to 4K tokens in one response.

Is Llama 3.1-70B open source?

Yes. Llama 3.1-70B's weights are downloadable from Hugging Face (meta-llama/Meta-Llama-3.1-70B-Instruct); check the license for commercial terms.

What are Llama 3.1-70B's strengths and weaknesses?

Relative to other ranked models, Llama 3.1-70B places best in reasoning, long context, agentic & tool use and lowest in math, coding, knowledge.

What is Llama 3.1-70B best at?

Its best category is agentic & tool use, where it ranks 112th on Noometry.