Long Context benchmark
CL-bench Life leaderboard
As of October 2026, GPT-5.5 has the highest published CL-bench Life score on Noometry at 22.2%, out of 13 models with results.
Last verified
About CL-bench Life
A description with primary sources is being prepared for this benchmark.
- Category
- Long Context
- Introduced
- 2026
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- epoch.ai
Top 13 models
- GPT-5.5 22.2%
- GPT-5.4 21.7%
- GPT-5.1 17.3%
- Claude Opus 4.6 17%
- Gemini 3.1 Pro Preview 16.9%
- DeepSeek V4 Pro 13.5%
- Kimi K2.5 13.2%
- Qwen3.5 Plus 12.4%
- Grok 4.20 (Non-Reasoning) 11.9%
- GLM-4.7 10.9%
- DeepSeek-V3.2-Exp 9.5%
- MiMo-V2-Pro 6.9%
- MiniMax-M2.5 6.3%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | GPT-5.5 | OpenAI | 22.2% | high | Epoch AI | |
| 2 | GPT-5.4 | OpenAI | 21.7% | xhigh | Epoch AI | |
| 3 | GPT-5.1 | OpenAI | 17.3% | high | Epoch AI | |
| 4 | Claude Opus 4.6 | Anthropic | 17% | high | Epoch AI | |
| 5 | Gemini 3.1 Pro Preview | 16.9% | Epoch AI | |||
| 6 | DeepSeek V4 Pro | 13.5% | high | Epoch AI | ||
| 7 | Kimi K2.5 | Moonshot AI | 13.2% | Epoch AI | ||
| 8 | Qwen3.5 Plus | 12.4% | Epoch AI | |||
| 9 | Grok 4.20 (Non-Reasoning) | xAI | 11.9% | Epoch AI | ||
| 10 | GLM-4.7 | Z.ai (Zhipu) | 10.9% | Epoch AI | ||
| 11 | DeepSeek-V3.2-Exp | 9.5% | thinking | Epoch AI | ||
| 12 | MiMo-V2-Pro | Xiaomi | 6.9% | Epoch AI | ||
| 13 | MiniMax-M2.5 | 6.3% | Epoch AI |
Compare the leaders
Other long context benchmarks
Frequently asked questions
Which model has the highest CL-bench Life score?
As of October 2026, GPT-5.5 has the highest published CL-bench Life score on Noometry at 22.2%, out of 13 models with results.
What is the best open-weight model on CL-bench Life?
DeepSeek V4 Pro has the highest CL-bench Life accuracy among open-weight models at 13.5%, ranking 6 of 13 overall.