Reasoning benchmark

# Bench to the Future 3 leaderboard

> Bench to the Future 3 results for 10 AI models, led by GLM-5.3 at 0.15. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/btf-3
- Last updated: 2026-10-10
- Title: Bench to the Future 3 Leaderboard (October 2026): Scores by Model

As of October 2026, GLM-5.3 has the highest published Bench to the Future 3 score on Noometry at 0.15, out of 10 models with results.

Last verified October 10, 2026

## About Bench to the Future 3

A description with primary sources is being prepared for this benchmark.

- **Category:** [Reasoning](https://noometry.com/best/reasoning)
- **Introduced:** 2026
- **Unit:** Raw score
- **Official site:** [epoch.ai](https://epoch.ai/benchmarks)

## Top 10 models

Top models on Bench to the Future 3

1.  GLM-5.3 0.15
2.  GLM-5.3-Flash 0.15
3.  GPT-5.5 0.14
4.  Claude Sonnet 5 0.14
5.  Muse Spark 1.3 0.14
6.  Claude Opus 4.8 0.14
7.  GPT-5.6 Sol 0.14
8.  GPT-6 Astra 0.14
9.  Claude Fable 5 0.13
10.  Claude Opus 5 0.12
11.  0.110.120.130.140.15

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

Bench to the Future 3 results by model
| # | Model | Provider | Rating | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GLM-5.3](https://noometry.com/models/glm-5-3) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 0.15 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [GLM-5.3-Flash](https://noometry.com/models/glm-5-3-flash) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 0.15 |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | [OpenAI](https://noometry.com/providers/openai) | 0.14 | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [Claude Sonnet 5](https://noometry.com/models/claude-sonnet-5) | [Anthropic](https://noometry.com/providers/anthropic) | 0.14 | xhigh | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [Muse Spark 1.3](https://noometry.com/models/muse-spark-1-3) |  [![](/logos/meta.svg) Meta](https://noometry.com/providers/meta) | 0.14 | xhigh | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [Claude Opus 4.8](https://noometry.com/models/claude-opus-4-8) | [Anthropic](https://noometry.com/providers/anthropic) | 0.14 | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [GPT-5.6 Sol](https://noometry.com/models/gpt-5-6-sol) | [OpenAI](https://noometry.com/providers/openai) | 0.14 | max | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [GPT-6 Astra](https://noometry.com/models/gpt-6-astra) | [OpenAI](https://noometry.com/providers/openai) | 0.14 | low | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [Claude Fable 5](https://noometry.com/models/claude-fable-5) | [Anthropic](https://noometry.com/providers/anthropic) | 0.13 | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 10 | [Claude Opus 5](https://noometry.com/models/claude-opus-5) | [Anthropic](https://noometry.com/providers/anthropic) | 0.12 | xhigh | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [GLM-5.3 vs GLM-5.3-Flash](https://noometry.com/compare/glm-5-3-vs-glm-5-3-flash)
-   [GLM-5.3 vs GPT-5.5](https://noometry.com/compare/glm-5-3-vs-gpt-5-5)
-   [GLM-5.3 vs Claude Sonnet 5](https://noometry.com/compare/claude-sonnet-5-vs-glm-5-3)
-   [GLM-5.3 vs Muse Spark 1.3](https://noometry.com/compare/glm-5-3-vs-muse-spark-1-3)
-   [GLM-5.3-Flash vs GPT-5.5](https://noometry.com/compare/glm-5-3-flash-vs-gpt-5-5)
-   [GLM-5.3-Flash vs Claude Sonnet 5](https://noometry.com/compare/claude-sonnet-5-vs-glm-5-3-flash)

## Other reasoning benchmarks

-   [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2)
-   [SimpleBench](https://noometry.com/benchmarks/simplebench)
-   [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning)
-   [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections)
-   [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1)
-   [CritPt](https://noometry.com/benchmarks/critpt)
-   [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles)
-   [EnigmaEval](https://noometry.com/benchmarks/enigmaeval)
-   [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization)
-   [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts)
-   [EBR-Bench](https://noometry.com/benchmarks/ebr-bench)
-   [LiveBench Reasoning](https://noometry.com/benchmarks/livebench-reasoning)

## Frequently asked questions

### Which model has the highest Bench to the Future 3 score?

As of October 2026, GLM-5.3 has the highest published Bench to the Future 3 score on Noometry at 0.15, out of 10 models with results.

### What is the best open-weight model on Bench to the Future 3?

GLM-5.3 has the highest Bench to the Future 3 score among open-weight models at 0.15, ranking 1 of 10 overall.

### Cite this page

Noometry. (2026). Bench to the Future 3 leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/btf-3

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/btf-3.md).
