Meta, open weights

Llama 3.1-405B

Llama 3.1-405B by Meta ranks 288th of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.7. Its strongest category is agentic & tool use, where it ranks 140th.

Last verified

Specifications

Noometry rank
#288 of 354
Index score
30.7
Evidence
Confirmed 42 results
Provider
Meta
Released
July 23, 2024
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
78 tokens/s Kagi
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Llama 3.1-405B category scores
  1. Coding 33.1
  2. Agentic & Tool Use 21.0
  3. Reasoning 16.8
  4. Math 18.4
  5. Knowledge 30.4
  6. Multilingual 40.7
  7. Instruction Following 65.9
  8. Long Context 38.4
  9. Writing & Preference 38.9
Llama 3.1-405B category ranks
CategoryScoreRankResults
Coding33.1#2622
Agentic & Tool Use21.0#1402
Reasoning16.8#3004
Math18.4#2904
Knowledge30.4#2275
Multilingual40.7#2141
Instruction Following65.9#2142
Long Context38.4#1971
Writing & Preference38.9#2515

Strengths and weaknesses

Categories where Llama 3.1-405B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 3.1-405B: strongest categories
CategoryScorevs medianRank
Long Context38.4−2.5#197 of 296, top 67%
Instruction Following65.9−5.4#214 of 305, top 71%
Multilingual40.7−6.7#214 of 297, top 73%

Weakest categories

Llama 3.1-405B: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use21.0−9.3#140 of 154, top 91%
Math18.4−18.2#290 of 327, top 89%
Reasoning16.8−6.8#300 of 350, top 86%

Closest competitors

The models ranked just above and below Llama 3.1-405B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 3.1-405B
ModelRankScoreBlended $/MSpeed
Gemma 3 27B#28430.8$0.1062Compare
Qwen1.5-72B#28530.8——Compare
Granite 3.0 2b Instruct#28630.8——Compare
Codellama 34b Instruct#28730.8——Compare
Yi-1.5-34B#28930.6——Compare
Codestral#29030.6$0.45271Compare
Llama-3.3-70B-Instruct#29130.6$0.16—Compare
GPT-4 Turbo#29230.5$15—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 3.1-405B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML21.4%#104 of 119, top 88%Epoch AI
LMArena Coding1283LMArena2026-10-08
LMArena Coding1291#204 of 294, top 70%LMArena2026-10-08

Agentic & Tool Use

Llama 3.1-405B Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
TheAgentCompany7.4%#9 of 14, top 65%Epoch AI
Cybench7.5%#19 of 21, top 91%Epoch AI

Reasoning

Llama 3.1-405B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench23%#68 of 77, top 89%Epoch AI
Kagi LLM Benchmark45%#70 of 99, top 71%Kagi LLM Benchmark
LMArena Hard Prompts1263LMArena2026-10-08
LMArena Hard Prompts1269#207 of 297, top 70%LMArena2026-10-08
DTBench61.4%#113 of 151, top 75%Epoch AI
BIG-Bench Hard82.9%#3 of 27, top 12%Epoch AI
Epoch Capabilities Index128.75#144 of 213, top 68%Epoch AI2024-07-23
ForecastBench59.9#35 of 72, top 49%Epoch AI
HellaSwag89.2%#3 of 29, top 11%Epoch AI
PIQA85.9%#4 of 27, top 15%Epoch AI
WinoGrande89.2%Best of 43Epoch AI
WinoGrande82.2%Epoch AI

Math

Llama 3.1-405B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20259.7%#133 of 173, top 77%Epoch AI2025-02-25
Omni-MATH24.9%#44 of 57, top 78%HELM Capabilities
LMArena Math1281#192 of 285, top 68%LMArena2026-10-08
LMArena Math1278LMArena2026-10-08
MATH Level 549.8%#46 of 79, top 59%Epoch AI2025-01-27

Knowledge

Llama 3.1-405B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond50.9%#127 of 186, top 69%Epoch AI2025-01-27
MMLU-Pro72.3%#33 of 58, top 57%HELM Capabilities
Confabulations (lower is better)17.6%#27 of 51, top 53%Lech Mazur benchmarks
GPQA (HELM)52.2%#30 of 57, top 53%HELM Capabilities
LMArena Expert1243#202 of 273, top 74%LMArena2026-10-08
LMArena Expert1229LMArena2026-10-08
ARC (AI2) Challenge95.3%#2 of 39, top 6%Epoch AI
MMLU84.5%#10 of 81, top 13%Epoch AI
MMLU84.4%Epoch AI
TriviaQA82.7%#8 of 25, top 32%Epoch AI

Multilingual

Llama 3.1-405B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1248#214 of 297, top 73%LMArena2026-10-08
LMArena Non-English1247LMArena2026-10-08
LMArena Chinese1242#213 of 285, top 75%LMArena2026-10-08
LMArena Chinese1234LMArena2026-10-08
LMArena French1271LMArena2026-10-08
LMArena French1279#172 of 223, top 78%LMArena2026-10-08
LMArena German1252#176 of 231, top 77%LMArena2026-10-08
LMArena German1251LMArena2026-10-08
LMArena Japanese1171LMArena2026-10-08
LMArena Japanese1208#153 of 211, top 73%LMArena2026-10-08
LMArena Korean1172LMArena2026-10-08
LMArena Korean1184#171 of 213, top 81%LMArena2026-10-08
LMArena Russian1265#199 of 283, top 71%LMArena2026-10-08
LMArena Russian1256LMArena2026-10-08
LMArena Spanish1260#179 of 226, top 80%LMArena2026-10-08
LMArena Spanish1253LMArena2026-10-08

Instruction Following

Llama 3.1-405B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval81.1%#38 of 57, top 67%HELM Capabilities
LMArena Instruction Following1259#202 of 298, top 68%LMArena2026-10-08
LMArena Instruction Following1259#202 of 298, top 68%LMArena2026-10-08

Long Context

Llama 3.1-405B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1260LMArena2026-10-08
LMArena Longer Query1266#211 of 291, top 73%LMArena2026-10-08

Writing & Preference

Llama 3.1-405B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1284#205 of 297, top 70%LMArena2026-10-08
LMArena Text1282LMArena2026-10-08
LMArena Creative Writing1260LMArena2026-10-08
LMArena Creative Writing1262#197 of 295, top 67%LMArena2026-10-08
EQ-Bench Creative Writing870#101 of 115, top 88%EQ-Bench
WildBench78.3%#40 of 57, top 71%HELM Capabilities
LMArena Multi-Turn1297#189 of 295, top 65%LMArena2026-10-08
LMArena Multi-Turn1287LMArena2026-10-08

Compare Llama 3.1-405B

Other Meta models

Frequently asked questions

How good is Llama 3.1-405B?

Llama 3.1-405B by Meta ranks 288th of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.7. Its strongest category is agentic & tool use, where it ranks 140th.

Is Llama 3.1-405B open source?

Yes. Llama 3.1-405B's weights are downloadable; check the license for commercial terms.

How fast is Llama 3.1-405B?

Llama 3.1-405B generated about 78 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Llama 3.1-405B's strengths and weaknesses?

Relative to other ranked models, Llama 3.1-405B places best in long context, instruction following, multilingual and lowest in agentic & tool use, math, reasoning.

What is Llama 3.1-405B best at?

Its best category is agentic & tool use, where it ranks 140th on Noometry.