Reasoning benchmark

CommonsenseQA 2.0 leaderboard

As of October 2026, GPT-3.5-turbo has the highest published CommonsenseQA 2.0 score on Noometry at 57%, out of 2 models with results.

Last verified

About CommonsenseQA 2.0

Yes/no commonsense questions collected through a model-in-the-loop game.

Category
Reasoning
Introduced
2022
Format
Yes/no
Unit
Percent (random guessing ≈ 50%)
Official site
allenai.github.io

Top 2 models

Top models on CommonsenseQA 2.0
  1. GPT-3.5-turbo 57%
  2. Llama 2-70B 50%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

CommonsenseQA 2.0 results by model
#ModelProviderScoreSettingSourceDate
1GPT-3.5-turbo OpenAI57%Epoch AI
2Llama 2-70B Meta50%Epoch AI

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does CommonsenseQA 2.0 measure?

Yes/no commonsense questions collected through a model-in-the-loop game.

Which model has the highest CommonsenseQA 2.0 score?

As of October 2026, GPT-3.5-turbo has the highest published CommonsenseQA 2.0 score on Noometry at 57%, out of 2 models with results.

What is the best open-weight model on CommonsenseQA 2.0?

Llama 2-70B has the highest CommonsenseQA 2.0 accuracy among open-weight models at 50%, ranking 2 of 2 overall.