Long Context benchmark
CL-bench leaderboard
As of October 2026, GPT-5.4 has the highest published CL-bench score on Noometry at 27.9%, out of 19 models with results.
Last verified
About CL-bench
A description with primary sources is being prepared for this benchmark.
- Category
- Long Context
- Introduced
- 2026
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- epoch.ai
Top 15 models
- GPT-5.4 27.9%
- GPT-5.1 23.7%
- Grok 4.20 (Non-Reasoning) 22.2%
- Claude Opus 4.5 21.1%
- Gemini 3.1 Pro Preview 20.8%
- Claude Opus 4.6 20.7%
- Qwen3.6 Plus 20.3%
- Qwen3.5 Plus 19.8%
- Kimi K2.5 19.3%
- GLM-5 18.7%
- GPT-5.2 18.2%
- o3 17.8%
- Kimi K2 (Jul 2025) 17.6%
- GLM-4.7 15.9%
- Gemini 3 Pro 15.8%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 | OpenAI | 27.9% | xhigh | Epoch AI | |
| 2 | GPT-5.1 | OpenAI | 23.7% | high | Epoch AI | |
| 3 | Grok 4.20 (Non-Reasoning) | xAI | 22.2% | Epoch AI | ||
| 4 | Claude Opus 4.5 | Anthropic | 21.1% | Epoch AI | ||
| 5 | Gemini 3.1 Pro Preview | 20.8% | Epoch AI | |||
| 6 | Claude Opus 4.6 | Anthropic | 20.7% | Epoch AI | ||
| 7 | Qwen3.6 Plus | 20.3% | Epoch AI | |||
| 8 | Qwen3.5 Plus | 19.8% | Epoch AI | |||
| 9 | Kimi K2.5 | Moonshot AI | 19.3% | Epoch AI | ||
| 10 | GLM-5 | Z.ai (Zhipu) | 18.7% | Epoch AI | ||
| 11 | GPT-5.2 | OpenAI | 18.2% | Epoch AI | ||
| 12 | o3 | OpenAI | 17.8% | high | Epoch AI | |
| 13 | Kimi K2 (Jul 2025) | Moonshot AI | 17.6% | Epoch AI | ||
| 14 | GLM-4.7 | Z.ai (Zhipu) | 15.9% | Epoch AI | ||
| 15 | Gemini 3 Pro | 15.8% | Epoch AI | |||
| 16 | MiMo-V2-Pro | Xiaomi | 15.7% | Epoch AI | ||
| 17 | Qwen3 Max | 14.5% | Epoch AI | |||
| 18 | DeepSeek-V3.2-Exp | 13.2% | thinking | Epoch AI | ||
| 19 | MiniMax-M2.5 | 11.4% | Epoch AI |
Compare the leaders
Other long context benchmarks
Frequently asked questions
Which model has the highest CL-bench score?
As of October 2026, GPT-5.4 has the highest published CL-bench score on Noometry at 27.9%, out of 19 models with results.
What is the best open-weight model on CL-bench?
Kimi K2.5 has the highest CL-bench accuracy among open-weight models at 19.3%, ranking 9 of 19 overall.