Coding benchmark
CadEval leaderboard
As of October 2026, o3 has the highest published CadEval score on Noometry at 74%, out of 14 models with results.
Last verified
About CadEval
Generating CAD code for 3D parts from text descriptions.
- Category
- Coding
- Introduced
- 2025
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- github.com
Top 14 models
- o3 74%
- Gemini 2.5 Pro 64%
- o4-mini 62%
- o1 56%
- Claude 3.7 Sonnet 54%
- o3-mini 54%
- Claude 3.5 Sonnet 48%
- GPT-4.1 42%
- Gemini 1.5 Pro (May 2024) 34%
- Claude 3.5 Haiku 32%
- Gemini 2.0 Flash (Feb 2025) 30%
- GPT-4o 26%
- GPT-4.1 mini 16%
- Claude 3 Haiku 12%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | o3 | OpenAI | 74% | medium | Epoch AI | |
| 2 | Gemini 2.5 Pro | 64% | Epoch AI | |||
| 3 | o4-mini | OpenAI | 62% | medium | Epoch AI | |
| 4 | o1 | OpenAI | 56% | medium | Epoch AI | |
| 5 | Claude 3.7 Sonnet | Anthropic | 54% | Epoch AI | ||
| 6 | o3-mini | OpenAI | 54% | medium | Epoch AI | |
| 7 | Claude 3.5 Sonnet | Anthropic | 48% | Epoch AI | ||
| 8 | GPT-4.1 | OpenAI | 42% | Epoch AI | ||
| 9 | Gemini 1.5 Pro (May 2024) | 34% | Epoch AI | |||
| 10 | Claude 3.5 Haiku | Anthropic | 32% | Epoch AI | ||
| 11 | Gemini 2.0 Flash (Feb 2025) | 30% | Epoch AI | |||
| 12 | GPT-4o | OpenAI | 26% | Epoch AI | ||
| 13 | GPT-4.1 mini | OpenAI | 16% | Epoch AI | ||
| 14 | Claude 3 Haiku | Anthropic | 12% | Epoch AI |
Compare the leaders
Frequently asked questions
What does CadEval measure?
Generating CAD code for 3D parts from text descriptions.
Which model has the highest CadEval score?
As of October 2026, o3 has the highest published CadEval score on Noometry at 74%, out of 14 models with results.