Reasoning benchmark

Thematic Generalization leaderboard

As of October 2026, Claude Opus 4.6 has the highest published Thematic Generalization score on Noometry at 80.6%, out of 23 models with results.

Last verified

About Thematic Generalization

Infer a narrow hidden theme from a few examples and anti-examples, then pick the one candidate that fits it among close distractors.

Category
Reasoning
Introduced
2025
Format
Pick the example
Unit
Percent (random guessing ≈ 0%)
Official site
github.com

Top 15 models

Top models on Thematic Generalization
  1. Claude Opus 4.6 80.6%
  2. GPT-5.4 80%
  3. Gemini 3.1 Pro Preview 79.4%
  4. Claude Sonnet 4.6 76.3%
  5. Claude Opus 4.7 72.8%
  6. GLM-5.1 69.8%
  7. Kimi K2.5 69.4%
  8. Qwen3.5 397B-A17B 65.1%
  9. DeepSeek-V3.2-Exp 65%
  10. Grok 4.20 (Non-Reasoning) 63.8%
  11. Gemini 3.1 Flash Lite 63.3%
  12. GPT-5.4 mini 61.7%
  13. Qwen3.6 Plus 59.5%
  14. Seed 2.0 Pro 57.1%
  15. Gemma 4 31B IT 53%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Compare the leaders

Other reasoning benchmarks

Frequently asked questions

What does Thematic Generalization measure?

Infer a narrow hidden theme from a few examples and anti-examples, then pick the one candidate that fits it among close distractors.

Which model has the highest Thematic Generalization score?

As of October 2026, Claude Opus 4.6 has the highest published Thematic Generalization score on Noometry at 80.6%, out of 23 models with results.

What is the best open-weight model on Thematic Generalization?

GLM-5.1 has the highest Thematic Generalization accuracy among open-weight models at 69.8%, ranking 6 of 23 overall.