Long Context benchmark
LMArena Longer Query leaderboard
As of October 2026, Gemini 4 Argon has the highest published LMArena Longer Query score on Noometry at 1549, out of 291 models with results.
Last verified
About LMArena Longer Query
LMArena ratings restricted to long prompts.
- Category
- Long Context
- Introduced
- 2024
- Format
- Pairwise human votes
- Unit
- Arena rating (Bradley–Terry)
- Official site
- lmarena.ai
Top 15 models
- Gemini 4 Argon 1549
- Claude Opus 5.5 1532
- Claude Fable 5.1 1522
- Claude Opus 4.6 1520
- Claude Opus 5 1515
- Claude Fable 5 1509
- Gemini 3.8 Flash 1508
- Claude Opus 4.7 1505
- MiMo-V2.6-Pro 1501
- Claude Sonnet 5.5 1498
- Kimi K3 1494
- Gemini 3.7 Flash 1492
- Qwen3.8 Max 1489
- Muse Spark 1.3 1488
- GPT-5.5 1484
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other long context benchmarks
Frequently asked questions
What does LMArena Longer Query measure?
LMArena ratings restricted to long prompts.
Which model has the highest LMArena Longer Query score?
As of October 2026, Gemini 4 Argon has the highest published LMArena Longer Query score on Noometry at 1549, out of 291 models with results.
What is the best open-weight model on LMArena Longer Query?
MiMo-V2.6-Pro has the highest LMArena Longer Query rating among open-weight models at 1501, ranking 9 of 291 overall.