Long Context benchmark
Fiction.LiveBench leaderboard
As of October 2026, GPT-5 has the highest published Fiction.LiveBench score on Noometry at 97.2%, out of 47 models with results.
Last verified
About Fiction.LiveBench
Comprehension questions about long fiction stories that require tracking plot and characters across the text. Scores here are at 16k tokens.
- Category
- Long Context
- Introduced
- 2025
- Format
- Long-context QA
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- fiction.live
Top 15 models
- GPT-5 97.2%
- o3-pro 97.2%
- Grok 4 94.4%
- Grok 4 Fast 94.4%
- Gemini 2.5 Pro 91.7%
- o3 88.9%
- Kimi K2.5 86.1%
- Claude 3.7 Sonnet 83.3%
- DeepSeek-V3.2-Exp 83.3%
- o1 83.3%
- QwQ-32B 83.3%
- Gemini 2.5 Flash 77.8%
- o4-mini 77.8%
- DeepSeek-R1 75%
- Qwen3 235B-A22B 75%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other long context benchmarks
Frequently asked questions
What does Fiction.LiveBench measure?
Comprehension questions about long fiction stories that require tracking plot and characters across the text. Scores here are at 16k tokens.
Which model has the highest Fiction.LiveBench score?
As of October 2026, GPT-5 has the highest published Fiction.LiveBench score on Noometry at 97.2%, out of 47 models with results.
What is the best open-weight model on Fiction.LiveBench?
Kimi K2.5 has the highest Fiction.LiveBench accuracy among open-weight models at 86.1%, ranking 7 of 47 overall.