Long Context benchmark

CL-bench Life leaderboard

As of October 2026, GPT-5.5 has the highest published CL-bench Life score on Noometry at 22.2%, out of 13 models with results.

Last verified

About CL-bench Life

A description with primary sources is being prepared for this benchmark.

Category
Long Context
Introduced
2026
Unit
Percent (random guessing ≈ 0%)
Official site
epoch.ai

Top 13 models

Top models on CL-bench Life
  1. GPT-5.5 22.2%
  2. GPT-5.4 21.7%
  3. GPT-5.1 17.3%
  4. Claude Opus 4.6 17%
  5. Gemini 3.1 Pro Preview 16.9%
  6. DeepSeek V4 Pro 13.5%
  7. Kimi K2.5 13.2%
  8. Qwen3.5 Plus 12.4%
  9. Grok 4.20 (Non-Reasoning) 11.9%
  10. GLM-4.7 10.9%
  11. DeepSeek-V3.2-Exp 9.5%
  12. MiMo-V2-Pro 6.9%
  13. MiniMax-M2.5 6.3%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other long context benchmarks

Frequently asked questions

Which model has the highest CL-bench Life score?

As of October 2026, GPT-5.5 has the highest published CL-bench Life score on Noometry at 22.2%, out of 13 models with results.

What is the best open-weight model on CL-bench Life?

DeepSeek V4 Pro has the highest CL-bench Life accuracy among open-weight models at 13.5%, ranking 6 of 13 overall.