Reasoning benchmark

HellaSwag leaderboard

As of October 2026, GPT-4 has the highest published HellaSwag score on Noometry at 95.3%, out of 29 models with results.

Last verified

About HellaSwag

Commonsense sentence completion with adversarially filtered wrong endings.

Category
Reasoning
Introduced
2019
Format
Multiple choice
Unit
Percent (random guessing ≈ 25%)
Official site
rowanzellers.com

Top 15 models

Top models on HellaSwag
  1. GPT-4 95.3%
  2. GPT-4 95.3%
  3. Llama 3.1-405B 89.2%
  4. Falcon-180B 89%
  5. DeepSeek-V3 88.9%
  6. DeepSeek-V2 (MoE-236B, May 2024) 87.1%
  7. Mixtral 8x7B 86.7%
  8. Llama 2-70B 85.3%
  9. Falcon-40B 85.3%
  10. Qwen2.5 72B Instruct 84.8%
  11. Qwen2.5-Coder-32B 83%
  12. Falcon 2 11B 82.9%
  13. Nemotron-4 15B 82.4%
  14. phi-3-medium 14B 82.4%
  15. Gemma 7B 82.2%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does HellaSwag measure?

Commonsense sentence completion with adversarially filtered wrong endings.

Which model has the highest HellaSwag score?

As of October 2026, GPT-4 has the highest published HellaSwag score on Noometry at 95.3%, out of 29 models with results.

What is the best open-weight model on HellaSwag?

Llama 3.1-405B has the highest HellaSwag accuracy among open-weight models at 89.2%, ranking 3 of 29 overall.