Long Context benchmark

# CL-bench Life leaderboard

> CL-bench Life results for 13 AI models, led by GPT-5.5 at 22.2%. What the benchmark measures, who runs it, and a source for every score.
- Canonical page: https://noometry.com/benchmarks/cl-bench-life
- Last updated: 2026-10-10
- Title: CL-bench Life Leaderboard (October 2026): Scores by Model

As of October 2026, GPT-5.5 has the highest published CL-bench Life score on Noometry at 22.2%, out of 13 models with results.

Last verified October 10, 2026

## About CL-bench Life

A description with primary sources is being prepared for this benchmark.

- **Category:** [Long Context](https://noometry.com/best/long-context)
- **Introduced:** 2026
- **Unit:** Percent (random guessing ≈ 0%)
- **Official site:** [epoch.ai](https://epoch.ai/benchmarks)

## Top 13 models

Top models on CL-bench Life

1.  GPT-5.5 22.2%
2.  GPT-5.4 21.7%
3.  GPT-5.1 17.3%
4.  Claude Opus 4.6 17%
5.  Gemini 3.1 Pro Preview 16.9%
6.  DeepSeek V4 Pro 13.5%
7.  Kimi K2.5 13.2%
8.  Qwen3.5 Plus 12.4%
9.  Grok 4.20 (Non-Reasoning) 11.9%
10.  GLM-4.7 10.9%
11.  DeepSeek-V3.2-Exp 9.5%
12.  MiMo-V2-Pro 6.9%
13.  MiniMax-M2.5 6.3%
14.  0102030

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## All results

CL-bench Life results by model
| # | Model | Provider | Score | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | [GPT-5.5](https://noometry.com/models/gpt-5-5) | [OpenAI](https://noometry.com/providers/openai) | 22.2% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 2 | [GPT-5.4](https://noometry.com/models/gpt-5-4) | [OpenAI](https://noometry.com/providers/openai) | 21.7% | xhigh | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 3 | [GPT-5.1](https://noometry.com/models/gpt-5-1) | [OpenAI](https://noometry.com/providers/openai) | 17.3% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 4 | [Claude Opus 4.6](https://noometry.com/models/claude-opus-4-6) | [Anthropic](https://noometry.com/providers/anthropic) | 17% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 5 | [Gemini 3.1 Pro Preview](https://noometry.com/models/gemini-3-1-pro-preview) |  [![](/logos/google.svg) Google](https://noometry.com/providers/google) | 16.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 6 | [DeepSeek V4 Pro](https://noometry.com/models/deepseek-v4-pro) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 13.5% | high | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 7 | [Kimi K2.5](https://noometry.com/models/kimi-k2-5) | [Moonshot AI](https://noometry.com/providers/moonshot) | 13.2% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 8 | [Qwen3.5 Plus](https://noometry.com/models/qwen3-5-plus) |  [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba) | 12.4% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 9 | [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20) | [xAI](https://noometry.com/providers/xai) | 11.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 10 | [GLM-4.7](https://noometry.com/models/glm-4-7) | [Z.ai (Zhipu)](https://noometry.com/providers/zai) | 10.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 11 | [DeepSeek-V3.2-Exp](https://noometry.com/models/deepseek-v3-2-exp) |  [![](/logos/deepseek.svg) DeepSeek](https://noometry.com/providers/deepseek) | 9.5% | thinking | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 12 | [MiMo-V2-Pro](https://noometry.com/models/mimo-v2-pro) | [Xiaomi](https://noometry.com/providers/xiaomi) | 6.9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| 13 | [MiniMax-M2.5](https://noometry.com/models/minimax-m2-5) |  [![](/logos/minimax.svg) MiniMax](https://noometry.com/providers/minimax) | 6.3% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

## Compare the leaders

-   [GPT-5.5 vs GPT-5.4](https://noometry.com/compare/gpt-5-4-vs-gpt-5-5)
-   [GPT-5.5 vs GPT-5.1](https://noometry.com/compare/gpt-5-1-vs-gpt-5-5)
-   [GPT-5.5 vs Claude Opus 4.6](https://noometry.com/compare/claude-opus-4-6-vs-gpt-5-5)
-   [GPT-5.5 vs Gemini 3.1 Pro Preview](https://noometry.com/compare/gemini-3-1-pro-preview-vs-gpt-5-5)
-   [GPT-5.4 vs GPT-5.1](https://noometry.com/compare/gpt-5-1-vs-gpt-5-4)
-   [GPT-5.4 vs Claude Opus 4.6](https://noometry.com/compare/claude-opus-4-6-vs-gpt-5-4)

## Other long context benchmarks

-   [Fiction.LiveBench](https://noometry.com/benchmarks/fiction-livebench)
-   [CL-bench](https://noometry.com/benchmarks/cl-bench)
-   [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query)

## Frequently asked questions

### Which model has the highest CL-bench Life score?

As of October 2026, GPT-5.5 has the highest published CL-bench Life score on Noometry at 22.2%, out of 13 models with results.

### What is the best open-weight model on CL-bench Life?

DeepSeek V4 Pro has the highest CL-bench Life accuracy among open-weight models at 13.5%, ranking 6 of 13 overall.

### Cite this page

Noometry. (2026). CL-bench Life leaderboard. Retrieved October 10, 2026, from https://noometry.com/benchmarks/cl-bench-life

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/benchmarks/cl-bench-life.md).
