Allen Institute for AI (Ai2), open weights

Tulu 3 (Tülu 3) 70B

Tulu 3 (Tülu 3) 70B by Allen Institute for AI (Ai2) ranks 251st of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.0. Its strongest category is reasoning, where it ranks 169th.

Last verified

Specifications

Noometry rank
#251 of 354
Index score
33.0
Evidence
Confirmed 14 results
Released
November 21, 2024
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Tulu 3 (Tülu 3) 70B category scores
  1. Coding 36.0
  2. Reasoning 23.9
  3. Math 14.2
  4. Knowledge 25.0
  5. Multilingual 39.9
  6. Instruction Following 64.8
  7. Long Context 37.1
  8. Writing & Preference 45.6
Tulu 3 (Tülu 3) 70B category ranks
CategoryScoreRankResults
Coding36.0#2141
Reasoning23.9#1691
Math14.2#3033
Knowledge25.0#2641
Multilingual39.9#2221
Instruction Following64.8#2271
Long Context37.1#2221
Writing & Preference45.6#2233

Strengths and weaknesses

Categories where Tulu 3 (Tülu 3) 70B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Tulu 3 (Tülu 3) 70B: strongest categories
CategoryScorevs medianRank
Reasoning23.9+0.3#169 of 350, top 49%
Coding36.0−2.7#214 of 340, top 63%
Writing & Preference45.6−8.2#223 of 312, top 72%

Weakest categories

Tulu 3 (Tülu 3) 70B: weakest categories
CategoryScorevs medianRank
Math14.2−22.4#303 of 327, top 93%
Knowledge25.0−12.3#264 of 314, top 85%
Long Context37.1−3.8#222 of 296, top 75%

Closest competitors

The models ranked just above and below Tulu 3 (Tülu 3) 70B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Tulu 3 (Tülu 3) 70B
ModelRankScoreBlended $/MSpeed
Granite 3.1 2b Instruct#24733.2——Compare
Gemma 2 2b IT#24833.1——Compare
Wizardlm 70b#24933.0——Compare
Phi 3 Medium 4k Instruct#25033.0——Compare
DeepSeek-R1-Distill-Qwen-14B#25232.7——Compare
Qwen1.5-14B#25332.7——Compare
Olmo 2 0325 32b Instruct#25432.7——Compare
gpt-oss-20b#25532.5$0.03696Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Tulu 3 (Tülu 3) 70B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1235#229 of 294, top 78%LMArena2026-10-08

Reasoning

Tulu 3 (Tülu 3) 70B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1220#229 of 297, top 78%LMArena2026-10-08

Math

Tulu 3 (Tülu 3) 70B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20254.4%#150 of 173, top 87%Epoch AI2025-03-07
LMArena Math1242#218 of 285, top 77%LMArena2026-10-08
MATH Level 542.7%#50 of 79, top 64%Epoch AI2025-01-27

Knowledge

Tulu 3 (Tülu 3) 70B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond46.3%#140 of 186, top 76%Epoch AI2025-01-27

Multilingual

Tulu 3 (Tülu 3) 70B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1236#222 of 297, top 75%LMArena2026-10-08
LMArena Chinese1249#208 of 285, top 73%LMArena2026-10-08
LMArena Russian1246#215 of 283, top 76%LMArena2026-10-08

Instruction Following

Tulu 3 (Tülu 3) 70B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1233#226 of 298, top 76%LMArena2026-10-08

Long Context

Tulu 3 (Tülu 3) 70B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1224#233 of 291, top 81%LMArena2026-10-08

Writing & Preference

Tulu 3 (Tülu 3) 70B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1256#224 of 297, top 76%LMArena2026-10-08
LMArena Creative Writing1231#221 of 295, top 75%LMArena2026-10-08
LMArena Multi-Turn1252#224 of 295, top 76%LMArena2026-10-08

Compare Tulu 3 (Tülu 3) 70B

Other Allen Institute for AI (Ai2) models

Frequently asked questions

How good is Tulu 3 (Tülu 3) 70B?

Tulu 3 (Tülu 3) 70B by Allen Institute for AI (Ai2) ranks 251st of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.0. Its strongest category is reasoning, where it ranks 169th.

Is Tulu 3 (Tülu 3) 70B open source?

Yes. Tulu 3 (Tülu 3) 70B's weights are downloadable; check the license for commercial terms.

What are Tulu 3 (Tülu 3) 70B's strengths and weaknesses?

Relative to other ranked models, Tulu 3 (Tülu 3) 70B places best in reasoning, coding, writing & preference and lowest in math, knowledge, long context.

What is Tulu 3 (Tülu 3) 70B best at?

Its best category is reasoning, where it ranks 169th on Noometry.