Instruction Following benchmark
IFEval leaderboard
As of October 2026, Grok-3 mini has the highest published IFEval score on Noometry at 95.1%, out of 57 models with results.
Last verified
About IFEval
Prompts with checkable instructions such as word counts, required keywords or output format, scored by strict automatic checks. HELM Capabilities run.
- Category
- Instruction Following
- Introduced
- 2023
- Size
- 541 prompts
- Format
- Verifiable constraints
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- crfm.stanford.edu
Top 15 models
- Grok-3 mini 95.1%
- Grok 4 94.9%
- GPT-5.1 93.5%
- GPT-5 Nano 93.2%
- o4-mini 92.8%
- GPT-5 Mini 92.7%
- Claude Opus 4 91.8%
- Llama 4 Maverick 90.8%
- GPT-4.1 mini 90.4%
- Gemini 2.5 Flash 89.8%
- Granite 4.0 H Small 89%
- Grok 3 88.4%
- Gemini 3 Pro 87.7%
- Mistral Large 87.7%
- GPT-5 87.5%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other instruction following benchmarks
Frequently asked questions
What does IFEval measure?
Prompts with checkable instructions such as word counts, required keywords or output format, scored by strict automatic checks. HELM Capabilities run.
Which model has the highest IFEval score?
As of October 2026, Grok-3 mini has the highest published IFEval score on Noometry at 95.1%, out of 57 models with results.
What is the best open-weight model on IFEval?
Llama 4 Maverick has the highest IFEval accuracy among open-weight models at 90.8%, ranking 8 of 57 overall.