Allen Institute for AI (Ai2), open weights

Llama 3.1 Tulu 3 8b

Llama 3.1 Tulu 3 8b by Allen Institute for AI (Ai2) ranks 224th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.7. Its strongest category is reasoning, where it ranks 188th.

Last verified

Specifications

Noometry rank
#224 of 354
Index score
35.7
Evidence
Confirmed 11 results
Released
Unknown
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Llama 3.1 Tulu 3 8b category scores
  1. Coding 34.4
  2. Reasoning 22.8
  3. Math 33.9
  4. Multilingual 35.4
  5. Instruction Following 61.3
  6. Long Context 35.8
  7. Writing & Preference 39.7
Llama 3.1 Tulu 3 8b category ranks
CategoryScoreRankResults
Coding34.4#2351
Reasoning22.8#1881
Math33.9#1981
Multilingual35.4#2461
Instruction Following61.3#2461
Long Context35.8#2391
Writing & Preference39.7#2453

Strengths and weaknesses

Categories where Llama 3.1 Tulu 3 8b places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 3.1 Tulu 3 8b: strongest categories
CategoryScorevs medianRank
Reasoning22.8−0.8#188 of 350, top 54%
Math33.9−2.7#198 of 327, top 61%
Coding34.4−4.3#235 of 340, top 70%

Weakest categories

Llama 3.1 Tulu 3 8b: weakest categories
CategoryScorevs medianRank
Multilingual35.4−12.0#246 of 297, top 83%
Long Context35.8−5.2#239 of 296, top 81%
Instruction Following61.3−9.9#246 of 305, top 81%

Closest competitors

The models ranked just above and below Llama 3.1 Tulu 3 8b. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 3.1 Tulu 3 8b
ModelRankScoreBlended $/MSpeed
Deepseek Coder v2#22035.9——Compare
C4ai Aya Expanse 32b#22135.9——Compare
Llama 3.1 Nemotron 51b Instruct#22235.9——Compare
Nemotron 4 340b Instruct#22335.9——Compare
Qwen3 14B#22535.5$0.6179Compare
DeepSeek-R1-Distill-Qwen-32B#22635.5——Compare
Magistral Medium#22735.2$2.750Compare
Gemini 2.0 Flash (Feb 2025)#22835.1—92Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 3.1 Tulu 3 8b Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1183#248 of 294, top 85%LMArena2026-10-08

Reasoning

Llama 3.1 Tulu 3 8b Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1174#246 of 297, top 83%LMArena2026-10-08

Math

Llama 3.1 Tulu 3 8b Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1195#233 of 285, top 82%LMArena2026-10-08

Multilingual

Llama 3.1 Tulu 3 8b Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1169#246 of 297, top 83%LMArena2026-10-08
LMArena Chinese1176#243 of 285, top 86%LMArena2026-10-08
LMArena Russian1193#238 of 283, top 85%LMArena2026-10-08

Instruction Following

Llama 3.1 Tulu 3 8b Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1174#245 of 298, top 83%LMArena2026-10-08

Long Context

Llama 3.1 Tulu 3 8b Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1181#248 of 291, top 86%LMArena2026-10-08

Writing & Preference

Llama 3.1 Tulu 3 8b Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1193#245 of 297, top 83%LMArena2026-10-08
LMArena Creative Writing1182#240 of 295, top 82%LMArena2026-10-08
LMArena Multi-Turn1154#251 of 295, top 86%LMArena2026-10-08

Compare Llama 3.1 Tulu 3 8b

Other Allen Institute for AI (Ai2) models

Frequently asked questions

How good is Llama 3.1 Tulu 3 8b?

Llama 3.1 Tulu 3 8b by Allen Institute for AI (Ai2) ranks 224th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.7. Its strongest category is reasoning, where it ranks 188th.

Is Llama 3.1 Tulu 3 8b open source?

Yes. Llama 3.1 Tulu 3 8b's weights are downloadable; check the license for commercial terms.

What are Llama 3.1 Tulu 3 8b's strengths and weaknesses?

Relative to other ranked models, Llama 3.1 Tulu 3 8b places best in reasoning, math, coding and lowest in multilingual, long context, instruction following.

What is Llama 3.1 Tulu 3 8b best at?

Its best category is reasoning, where it ranks 188th on Noometry.