Allen Institute for AI (Ai2), open weights
Olmo 3.1 32b Think
Olmo 3.1 32b Think by Allen Institute for AI (Ai2) ranks 191st of 354 ranked models on the Noometry Index as of October 2026, with a score of 37.9. Its strongest category is reasoning, where it ranks 150th.
Last verified
Specifications
- Noometry rank
- #191 of 354
- Index score
- 37.9
- Evidence
- Confirmed 15 results
- Provider
Allen Institute for AI (Ai2)
- Released
- Unknown
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 37.7
- Reasoning 25.2
- Math 36.3
- Knowledge 35.7
- Multilingual 38.1
- Instruction Following 65.6
- Long Context 38.6
- Writing & Preference 46.2
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 37.7 | #189 | 1 |
| Reasoning | 25.2 | #150 | 1 |
| Math | 36.3 | #168 | 1 |
| Knowledge | 35.7 | #181 | 1 |
| Multilingual | 38.1 | #231 | 1 |
| Instruction Following | 65.6 | #218 | 1 |
| Long Context | 38.6 | #195 | 1 |
| Writing & Preference | 46.2 | #220 | 3 |
Strengths and weaknesses
Categories where Olmo 3.1 32b Think places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multilingual | 38.1 | −9.3 | #231 of 297, top 78% |
| Instruction Following | 65.6 | −5.7 | #218 of 305, top 72% |
| Writing & Preference | 46.2 | −7.5 | #220 of 312, top 71% |
Closest competitors
The models ranked just above and below Olmo 3.1 32b Think. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Sonar | #187 | 38.5 | $1 | — | Compare |
| MiniMax-M2.5 | #188 | 38.3 | $0.52 | 49 | Compare |
| Nova Premier 1.0 | #189 | 38.3 | $5 | 9 | Compare |
| Qwen3-Coder 480B-A35B Instruct | #190 | 38.1 | $3 | 67 | Compare |
| GPT-5-Codex | #192 | 37.9 | $3.44 | 95 | Compare |
| Hunyuan Standard 2025 02 10 | #193 | 37.9 | — | — | Compare |
| Gemini 2.0 Flash-Lite | #194 | 37.8 | — | — | Compare |
| DeepSeek-R1-Distill-Llama-70B | #195 | 37.8 | — | 18 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 1291 | #203 of 294, top 70% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Hard Prompts | 1272 | #203 of 297, top 69% | LMArena | 2026-10-08 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 1305 | #181 of 285, top 64% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Expert | 1295 | #179 of 273, top 66% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1209 | #231 of 297, top 78% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1242 | #214 of 285, top 76% | LMArena | 2026-10-08 | |
| LMArena French | 1260 | #180 of 223, top 81% | LMArena | 2026-10-08 | |
| LMArena German | 1262 | #168 of 231, top 73% | LMArena | 2026-10-08 | |
| LMArena Russian | 1193 | #237 of 283, top 84% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1289 | #166 of 226, top 74% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1247 | #216 of 298, top 73% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1272 | #209 of 291, top 72% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1272 | #216 of 297, top 73% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1226 | #222 of 295, top 76% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1252 | #223 of 295, top 76% | LMArena | 2026-10-08 |
Compare Olmo 3.1 32b Think
- Olmo 3.1 32b Think vs Qwen3-Coder 480B-A35B Instruct
- Olmo 3.1 32b Think vs GPT-5-Codex
- Olmo 3.1 32b Think vs Nova Premier 1.0
- Olmo 3.1 32b Think vs Hunyuan Standard 2025 02 10
- Olmo 3.1 32b Think vs MiniMax-M2.5
- Olmo 3.1 32b Think vs Gemini 2.0 Flash-Lite
- Olmo 3.1 32b Think vs GPT-6 Astra
- Olmo 3.1 32b Think vs Claude Fable 5.1
- Olmo 3.1 32b Think vs Gemini 3.8 Flash
- Olmo 3.1 32b Think vs Kimi K3
- Olmo 3.1 32b Think vs Grok 4.6
- Olmo 3.1 32b Think vs Qwen3.8 Max
- Olmo 3.1 32b Think vs GLM-5.3
- Olmo 3.1 32b Think vs Muse Spark 1.3
Other Allen Institute for AI (Ai2) models
Frequently asked questions
How good is Olmo 3.1 32b Think?
Olmo 3.1 32b Think by Allen Institute for AI (Ai2) ranks 191st of 354 ranked models on the Noometry Index as of October 2026, with a score of 37.9. Its strongest category is reasoning, where it ranks 150th.
Is Olmo 3.1 32b Think open source?
Yes. Olmo 3.1 32b Think's weights are downloadable; check the license for commercial terms.
What are Olmo 3.1 32b Think's strengths and weaknesses?
Relative to other ranked models, Olmo 3.1 32b Think places best in reasoning, math, coding and lowest in multilingual, instruction following, writing & preference.
What is Olmo 3.1 32b Think best at?
Its best category is reasoning, where it ranks 150th on Noometry.