Meta, open weights

Llama 3-8B

Llama 3-8B by Meta ranks 344th of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is long context, where it ranks 251st.

Last verified

Specifications

Noometry rank
#344 of 354
Index score
25.5
Evidence
Confirmed 34 results
Provider
Meta
Released
April 18, 2024
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Llama 3-8B category scores
  1. Coding 31.0
  2. Reasoning 14.3
  3. Math 8.8
  4. Knowledge 7.8
  5. Multilingual 30.8
  6. Instruction Following 58.4
  7. Long Context 34.2
  8. Writing & Preference 37.5
Llama 3-8B category ranks
CategoryScoreRankResults
Coding31.0#2893
Reasoning14.3#3263
Math8.8#3233
Knowledge7.8#3082
Multilingual30.8#2611
Instruction Following58.4#2601
Long Context34.2#2511
Writing & Preference37.5#2563

Strengths and weaknesses

Categories where Llama 3-8B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 3-8B: strongest categories
CategoryScorevs medianRank
Writing & Preference37.5−16.3#256 of 312, top 83%
Long Context34.2−6.8#251 of 296, top 85%
Coding31.0−7.7#289 of 340, top 85%

Weakest categories

Llama 3-8B: weakest categories
CategoryScorevs medianRank
Math8.8−27.7#323 of 327, top 99%
Knowledge7.8−29.6#308 of 314, top 99%
Reasoning14.3−9.3#326 of 350, top 94%

Closest competitors

The models ranked just above and below Llama 3-8B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 3-8B
ModelRankScoreBlended $/MSpeed
Claude 3 Haiku#34025.9—41Compare
Gemma 2 9B#34125.9——Compare
Dolly 2.0-12b#34225.5——Compare
GPT-4o mini#34325.5$0.26120Compare
Claude 2.1#34525.2——Compare
Claude 2#34625.0——Compare
DeepSeek LLM 67B#34724.9——Compare
Llama 13b#34824.4——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 3-8B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
BigCodeBench Instruct31.9%#54 of 64, top 85%BigCodeBench2024-04-18
LMArena Coding1152#257 of 294, top 88%LMArena2026-10-08
BigCodeBench Complete28.8%BigCodeBench2024-04-18
BigCodeBench Complete36.9%#58 of 66, top 88%BigCodeBench2024-04-18
HumanEval+56.7%#32 of 45, top 72%EvalPlus
HumanEval+29.3%EvalPlus
MBPP+51.6%EvalPlus
MBPP+54.8%#30 of 38, top 79%EvalPlus

Reasoning

Llama 3-8B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Chess Puzzles0%#122 of 129, top 95%Epoch AI2026-08-28
LMArena Hard Prompts1133#258 of 297, top 87%LMArena2026-10-08
DTBench43.9%#148 of 151, top 99%Epoch AI
Adversarial NLI57.3%#3 of 9, top 34%Epoch AI
Epoch Capabilities Index116.45#183 of 213, top 86%Epoch AI2024-04-18
ForecastBench58.6#46 of 72, top 64%Epoch AI
ForecastBench52.9Epoch AI
WinoGrande75.7%#22 of 43, top 52%Epoch AI
WinoGrande65%Epoch AI

Math

Llama 3-8B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20251.9%#162 of 173, top 94%Epoch AI2026-08-30
LMArena Math1151#251 of 285, top 89%LMArena2026-10-08
MATH Level 56.1%#76 of 79, top 97%Epoch AI2025-01-27

Knowledge

Llama 3-8B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond26.1%#179 of 186, top 97%Epoch AI2025-01-27
LMArena Expert1113#247 of 273, top 91%LMArena2026-10-08
ARC (AI2) Challenge82.8%#12 of 39, top 31%Epoch AI
MMLU68.8%#48 of 81, top 60%Epoch AI
OpenBookQA82.6%#6 of 19, top 32%Epoch AI
TriviaQA67.7%#20 of 25, top 80%Epoch AI

Multilingual

Llama 3-8B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1098#261 of 297, top 88%LMArena2026-10-08
LMArena Chinese1076#259 of 285, top 91%LMArena2026-10-08
LMArena French1159#204 of 223, top 92%LMArena2026-10-08
LMArena German1104#210 of 231, top 91%LMArena2026-10-08
LMArena Japanese967#203 of 211, top 97%LMArena2026-10-08
LMArena Korean1004#201 of 213, top 95%LMArena2026-10-08
LMArena Russian1109#256 of 283, top 91%LMArena2026-10-08
LMArena Spanish1173#199 of 226, top 89%LMArena2026-10-08

Instruction Following

Llama 3-8B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1127#258 of 298, top 87%LMArena2026-10-08

Long Context

Llama 3-8B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1128#260 of 291, top 90%LMArena2026-10-08

Writing & Preference

Llama 3-8B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1166#254 of 297, top 86%LMArena2026-10-08
LMArena Creative Writing1150#250 of 295, top 85%LMArena2026-10-08
LMArena Multi-Turn1152#253 of 295, top 86%LMArena2026-10-08

Compare Llama 3-8B

Other Meta models

Frequently asked questions

How good is Llama 3-8B?

Llama 3-8B by Meta ranks 344th of 354 ranked models on the Noometry Index as of October 2026, with a score of 25.5. Its strongest category is long context, where it ranks 251st.

Is Llama 3-8B open source?

Yes. Llama 3-8B's weights are downloadable; check the license for commercial terms.

What are Llama 3-8B's strengths and weaknesses?

Relative to other ranked models, Llama 3-8B places best in writing & preference, long context, coding and lowest in math, knowledge, reasoning.

What is Llama 3-8B best at?

Its best category is long context, where it ranks 251st on Noometry.