Reasoning benchmark
CommonsenseQA 2.0 leaderboard
As of October 2026, GPT-3.5-turbo has the highest published CommonsenseQA 2.0 score on Noometry at 57%, out of 2 models with results.
Last verified
About CommonsenseQA 2.0
Yes/no commonsense questions collected through a model-in-the-loop game.
- Category
- Reasoning
- Introduced
- 2022
- Format
- Yes/no
- Unit
- Percent (random guessing ≈ 50%)
- Official site
- allenai.github.io
Top 2 models
- GPT-3.5-turbo 57%
- Llama 2-70B 50%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | GPT-3.5-turbo | OpenAI | 57% | Epoch AI | ||
| 2 | Llama 2-70B | 50% | Epoch AI |
Compare the leaders
Frequently asked questions
What does CommonsenseQA 2.0 measure?
Yes/no commonsense questions collected through a model-in-the-loop game.
Which model has the highest CommonsenseQA 2.0 score?
As of October 2026, GPT-3.5-turbo has the highest published CommonsenseQA 2.0 score on Noometry at 57%, out of 2 models with results.
What is the best open-weight model on CommonsenseQA 2.0?
Llama 2-70B has the highest CommonsenseQA 2.0 accuracy among open-weight models at 50%, ranking 2 of 2 overall.