Allen Institute for AI (Ai2), open weights

Olmo 3.1 32b Think

Olmo 3.1 32b Think by Allen Institute for AI (Ai2) ranks 191st of 354 ranked models on the Noometry Index as of October 2026, with a score of 37.9. Its strongest category is reasoning, where it ranks 150th.

Last verified

Specifications

Noometry rank
#191 of 354
Index score
37.9
Evidence
Confirmed 15 results
Released
Unknown
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Olmo 3.1 32b Think category scores
  1. Coding 37.7
  2. Reasoning 25.2
  3. Math 36.3
  4. Knowledge 35.7
  5. Multilingual 38.1
  6. Instruction Following 65.6
  7. Long Context 38.6
  8. Writing & Preference 46.2
Olmo 3.1 32b Think category ranks
CategoryScoreRankResults
Coding37.7#1891
Reasoning25.2#1501
Math36.3#1681
Knowledge35.7#1811
Multilingual38.1#2311
Instruction Following65.6#2181
Long Context38.6#1951
Writing & Preference46.2#2203

Strengths and weaknesses

Categories where Olmo 3.1 32b Think places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Olmo 3.1 32b Think: strongest categories
CategoryScorevs medianRank
Reasoning25.2+1.6#150 of 350, top 43%
Math36.3−0.3#168 of 327, top 52%
Coding37.7−1.0#189 of 340, top 56%

Weakest categories

Olmo 3.1 32b Think: weakest categories
CategoryScorevs medianRank
Multilingual38.1−9.3#231 of 297, top 78%
Instruction Following65.6−5.7#218 of 305, top 72%
Writing & Preference46.2−7.5#220 of 312, top 71%

Closest competitors

The models ranked just above and below Olmo 3.1 32b Think. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Olmo 3.1 32b Think
ModelRankScoreBlended $/MSpeed
Sonar#18738.5$1—Compare
MiniMax-M2.5#18838.3$0.5249Compare
Nova Premier 1.0#18938.3$59Compare
Qwen3-Coder 480B-A35B Instruct#19038.1$367Compare
GPT-5-Codex#19237.9$3.4495Compare
Hunyuan Standard 2025 02 10#19337.9——Compare
Gemini 2.0 Flash-Lite#19437.8——Compare
DeepSeek-R1-Distill-Llama-70B#19537.8—18Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Olmo 3.1 32b Think Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1291#203 of 294, top 70%LMArena2026-10-08

Reasoning

Olmo 3.1 32b Think Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1272#203 of 297, top 69%LMArena2026-10-08

Math

Olmo 3.1 32b Think Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1305#181 of 285, top 64%LMArena2026-10-08

Knowledge

Olmo 3.1 32b Think Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1295#179 of 273, top 66%LMArena2026-10-08

Multilingual

Olmo 3.1 32b Think Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1209#231 of 297, top 78%LMArena2026-10-08
LMArena Chinese1242#214 of 285, top 76%LMArena2026-10-08
LMArena French1260#180 of 223, top 81%LMArena2026-10-08
LMArena German1262#168 of 231, top 73%LMArena2026-10-08
LMArena Russian1193#237 of 283, top 84%LMArena2026-10-08
LMArena Spanish1289#166 of 226, top 74%LMArena2026-10-08

Instruction Following

Olmo 3.1 32b Think Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1247#216 of 298, top 73%LMArena2026-10-08

Long Context

Olmo 3.1 32b Think Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1272#209 of 291, top 72%LMArena2026-10-08

Writing & Preference

Olmo 3.1 32b Think Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1272#216 of 297, top 73%LMArena2026-10-08
LMArena Creative Writing1226#222 of 295, top 76%LMArena2026-10-08
LMArena Multi-Turn1252#223 of 295, top 76%LMArena2026-10-08

Compare Olmo 3.1 32b Think

Other Allen Institute for AI (Ai2) models

Frequently asked questions

How good is Olmo 3.1 32b Think?

Olmo 3.1 32b Think by Allen Institute for AI (Ai2) ranks 191st of 354 ranked models on the Noometry Index as of October 2026, with a score of 37.9. Its strongest category is reasoning, where it ranks 150th.

Is Olmo 3.1 32b Think open source?

Yes. Olmo 3.1 32b Think's weights are downloadable; check the license for commercial terms.

What are Olmo 3.1 32b Think's strengths and weaknesses?

Relative to other ranked models, Olmo 3.1 32b Think places best in reasoning, math, coding and lowest in multilingual, instruction following, writing & preference.

What is Olmo 3.1 32b Think best at?

Its best category is reasoning, where it ranks 150th on Noometry.