Reasoning benchmark

# ForecastBench leaderboard

> ForecastBench results for 72 AI models, led by o3 at 62.5. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/forecastbench
- Last updated: 2026-10-10
- Title: ForecastBench Leaderboard (October 2026): Scores by Model

As of October 2026, o3 has the highest published ForecastBench score on Noometry at 62.5, out of 72 models with results.

Last verified October 10, 2026

## About ForecastBench

Forecasting real-world events that resolve after the model's training cutoff.

- **Category:** [Reasoning](https://noometry.com/best/reasoning)
- **Introduced:** 2024
- **Format:** Probabilistic forecasts
- **Unit:** Raw score
- **Official site:** [www.forecastbench.org](https://www.forecastbench.org)

## Top 15 models

Top models on ForecastBench

1.  o3 62.5
2.  Claude Opus 4.1 62
3.  Claude Sonnet 4.6 62
4.  Claude Sonnet 4.5 61.9
5.  Claude 3.7 Sonnet 61.8
6.  o4-mini 61.8
7.  GPT-4.5 61.7
8.  GPT-4.1 61.5
9.  Claude Haiku 4.5 61.4
10.  GPT-5 61.4
11.  Grok 4.20 (Non-Reasoning) 61.4
12.  MiniMax-M3 61.4
13.  Gemini 2.5 Pro 61.3
14.  Gemini 3 Pro 61.2
15.  Claude Opus 4 61.1
16.  60.561.061.562.062.5

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

ForecastBench results by model
| # | Model | Provider | Rating | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [o3](https://noometry.com/models/o3) | [OpenAI](https://noometry.com/providers/openai) | 62.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [Claude Opus 4.1](https://noometry.com/models/claude-opus-4-1) | [Anthropic](https://noometry.com/providers/anthropic) | 62 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [Claude Sonnet 4.6](https://noometry.com/models/claude-sonnet-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 62 | 16K | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [Claude Sonnet 4.5](https://noometry.com/models/claude-sonnet-4-5) | [Anthropic](https://noometry.com/providers/anthropic) | 61.9 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [Claude 3.7 Sonnet](https://noometry.com/models/claude-3-7-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 61.8 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [o4-mini](https://noometry.com/models/o4-mini) | [OpenAI](https://noometry.com/providers/openai) | 61.8 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [GPT-4.5](https://noometry.com/models/gpt-4-5) | [OpenAI](https://noometry.com/providers/openai) | 61.7 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [GPT-4.1](https://noometry.com/models/gpt-4-1) | [OpenAI](https://noometry.com/providers/openai) | 61.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [Claude Haiku 4.5](https://noometry.com/models/claude-haiku-4-5) | [Anthropic](https://noometry.com/providers/anthropic) | 61.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 10 | [GPT-5](https://noometry.com/models/gpt-5) | [OpenAI](https://noometry.com/providers/openai) | 61.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 11 | [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20) | [xAI](https://noometry.com/providers/xai) | 61.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 12 | [MiniMax-M3](https://noometry.com/models/minimax-m3) |  [![](/logos/minimax.svg) MiniMax](https://noometry.com/providers/minimax) | 61.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 13 | [Gemini 2.5 Pro](https://noometry.com/models/gemini-2-5-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 61.3 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 14 | [Gemini 3 Pro](https://noometry.com/models/gemini-3-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 61.2 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 15 | [Claude Opus 4](https://noometry.com/models/claude-opus-4) | [Anthropic](https://noometry.com/providers/anthropic) | 61.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 16 | [Claude Sonnet 5](https://noometry.com/models/claude-sonnet-5) | [Anthropic](https://noometry.com/providers/anthropic) | 61.1 | 16K | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 17 | [Kimi K3](https://noometry.com/models/kimi-k3) | [Moonshot AI](https://noometry.com/providers/moonshot) | 61.1 | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 18 | [GLM-5](https://noometry.com/models/glm-5) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 61 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 19 | [GPT-5 Mini](https://noometry.com/models/gpt-5-mini) | [OpenAI](https://noometry.com/providers/openai) | 61 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 20 | [Grok 4.1 Fast](https://noometry.com/models/grok-4-1-fast) | [xAI](https://noometry.com/providers/xai) | 61 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 21 | [Grok 4](https://noometry.com/models/grok-4) | [xAI](https://noometry.com/providers/xai) | 60.9 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 22 | [Claude 3.5 Sonnet](https://noometry.com/models/claude-3-5-sonnet) | [Anthropic](https://noometry.com/providers/anthropic) | 60.7 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 23 | [Claude Opus 4.5](https://noometry.com/models/claude-opus-4-5) | [Anthropic](https://noometry.com/providers/anthropic) | 60.7 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 24 | [Gemini 2.5 Flash](https://noometry.com/models/gemini-2-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 60.6 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 25 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | [OpenAI](https://noometry.com/providers/openai) | 60.6 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 26 | [Grok 4 Fast](https://noometry.com/models/grok-4-fast) | [xAI](https://noometry.com/providers/xai) | 60.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 27 | [Claude Opus 4.7](https://noometry.com/models/claude-opus-4-7) | [Anthropic](https://noometry.com/providers/anthropic) | 60.3 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 28 | [Grok 4.3](https://noometry.com/models/grok-4-3) | [xAI](https://noometry.com/providers/xai) | 60.3 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 29 | [Claude Sonnet 4](https://noometry.com/models/claude-sonnet-4) | [Anthropic](https://noometry.com/providers/anthropic) | 60.2 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 30 | [Kimi K2 (Jul 2025)](https://noometry.com/models/kimi-k2) | [Moonshot AI](https://noometry.com/providers/moonshot) | 60.2 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 31 | [GPT-5.2](https://noometry.com/models/gpt-5-2) | [OpenAI](https://noometry.com/providers/openai) | 60.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 32 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 60 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 33 | [DeepSeek-R1](https://noometry.com/models/deepseek-r1) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 60 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 34 | [Claude Opus 4.8](https://noometry.com/models/claude-opus-4-8) | [Anthropic](https://noometry.com/providers/anthropic) | 59.9 | 24K | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 35 | [Llama 3.1-405B](https://noometry.com/models/llama-3-1-405b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 59.9 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 36 | [Qwen3 235B-A22B](https://noometry.com/models/qwen3-235b-a22b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 59.7 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 37 | [o3-mini](https://noometry.com/models/o3-mini) | [OpenAI](https://noometry.com/providers/openai) | 59.6 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 38 | [GPT-5.4](https://noometry.com/models/gpt-5-4) | [OpenAI](https://noometry.com/providers/openai) | 59.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 39 | [GPT-4 Turbo](https://noometry.com/models/gpt-4-turbo) | [OpenAI](https://noometry.com/providers/openai) | 59.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 40 | [GLM-4.5-Air](https://noometry.com/models/glm-4-5-air) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 59.2 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 41 | [DeepSeek-V3](https://noometry.com/models/deepseek-v3) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 59.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 42 | [GPT-5 Nano](https://noometry.com/models/gpt-5-nano) | [OpenAI](https://noometry.com/providers/openai) | 59.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 43 | [Gemini 3.1 Pro Preview](https://noometry.com/models/gemini-3-1-pro-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 59 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 44 | [Gemini 3.5 Flash](https://noometry.com/models/gemini-3-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 59 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 45 | [Llama-3.3-70B-Instruct](https://noometry.com/models/llama-3-3-70b-instruct) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 58.6 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 46 | [Llama 3-8B](https://noometry.com/models/llama-3-8b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 58.6 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 47 | [Gemini 3 Flash Preview](https://noometry.com/models/gemini-3-flash-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 58.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 48 | [Claude 3 Opus](https://noometry.com/models/claude-3-opus) | [Anthropic](https://noometry.com/providers/anthropic) | 58.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 49 | [Gemini 1.5 Pro (May 2024)](https://noometry.com/models/gemini-1-5-pro) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 58.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 50 | [QwQ-32B](https://noometry.com/models/qwq-32b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 58.3 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 51 | [GPT-5.1](https://noometry.com/models/gpt-5-1) | [OpenAI](https://noometry.com/providers/openai) | 58.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 52 | [DeepSeek-V3.1](https://noometry.com/models/deepseek-v3-1) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 58 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 53 | [GPT-4](https://noometry.com/models/gpt-4) | [OpenAI](https://noometry.com/providers/openai) | 57.8 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 54 | [GPT-4o](https://noometry.com/models/gpt-4o) | [OpenAI](https://noometry.com/providers/openai) | 57.7 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 55 | [Qwen1.5-110B](https://noometry.com/models/qwen1-5-110b) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 57.7 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 56 | [Llama 4 Maverick](https://noometry.com/models/llama-4-maverick) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 57.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 57 | [Llama 4 Scout](https://noometry.com/models/llama-4-scout) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 57.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 58 | [Qwen2.5 72B Instruct](https://noometry.com/models/qwen2-5-72b-instruct) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 57.5 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 59 | [GPT-5.4 nano](https://noometry.com/models/gpt-5-4-nano) | [OpenAI](https://noometry.com/providers/openai) | 57.3 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 60 | [Gemini 2.0 Flash-Lite](https://noometry.com/models/gemini-2-0-flash-lite) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 57.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 61 | [Llama 3-70B](https://noometry.com/models/llama-3-70b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 57.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 62 | [Mistral Large](https://noometry.com/models/mistral-large) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 57.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 63 | [GPT-5.4 mini](https://noometry.com/models/gpt-5-4-mini) | [OpenAI](https://noometry.com/providers/openai) | 57 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 64 | [Mixtral 8x22B](https://noometry.com/models/mixtral-8x22b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 56.3 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 65 | [Mixtral 8x7B](https://noometry.com/models/mixtral-8x7b) |  [![](/logos/mistral.svg) Mistral AI](https://noometry.com/providers/mistral) | 56.3 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 66 | [DeepSeek V4 Pro](https://noometry.com/models/deepseek-v4-pro) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 56.1 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 67 | [Gemini 3.1 Flash Lite](https://noometry.com/models/gemini-3-1-flash-lite) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 54.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 68 | [Claude 2.1](https://noometry.com/models/claude-2-1) | [Anthropic](https://noometry.com/providers/anthropic) | 54.2 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 69 | [Gemini 1.5 Flash (May 2024)](https://noometry.com/models/gemini-1-5-flash) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 53.9 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 70 | [Claude 3 Haiku](https://noometry.com/models/claude-3-haiku) | [Anthropic](https://noometry.com/providers/anthropic) | 53.2 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 71 | [Llama 2-70B](https://noometry.com/models/llama-2-70b) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 51.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 72 | [GPT-3.5-turbo](https://noometry.com/models/gpt-3-5-turbo) | [OpenAI](https://noometry.com/providers/openai) | 50.4 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [o3 vs Claude Opus 4.1](https://noometry.com/compare/claude-opus-4-1-vs-o3)
-   [o3 vs Claude Sonnet 4.6](https://noometry.com/compare/claude-sonnet-4-6-vs-o3)
-   [o3 vs Claude Sonnet 4.5](https://noometry.com/compare/claude-sonnet-4-5-vs-o3)
-   [o3 vs Claude 3.7 Sonnet](https://noometry.com/compare/claude-3-7-sonnet-vs-o3)
-   [Claude Opus 4.1 vs Claude Sonnet 4.6](https://noometry.com/compare/claude-opus-4-1-vs-claude-sonnet-4-6)
-   [Claude Opus 4.1 vs Claude Sonnet 4.5](https://noometry.com/compare/claude-opus-4-1-vs-claude-sonnet-4-5)

## Other reasoning benchmarks

-   [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2)
-   [SimpleBench](https://noometry.com/benchmarks/simplebench)
-   [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning)
-   [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections)
-   [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1)
-   [CritPt](https://noometry.com/benchmarks/critpt)
-   [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles)
-   [EnigmaEval](https://noometry.com/benchmarks/enigmaeval)
-   [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization)
-   [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts)
-   [EBR-Bench](https://noometry.com/benchmarks/ebr-bench)
-   [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning)

## Frequently asked questions

### What does ForecastBench measure?

Forecasting real-world events that resolve after the model's training cutoff.

### Which model has the highest ForecastBench score?

As of October 2026, o3 has the highest published ForecastBench score on Noometry at 62.5, out of 72 models with results.

### What is the best open-weight model on ForecastBench?

MiniMax-M3 has the highest ForecastBench score among open-weight models at 61.4, ranking 12 of 72 overall.

### Cite this page

Noometry. (2026). ForecastBench leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/forecastbench

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/forecastbench.md).
