Instruction Following benchmark

LiveBench Instruction Following leaderboard

As of October 2026, GPT-5.1 has the highest published LiveBench Instruction Following score on Noometry at 93.3%, out of 39 models with results.

Last verified

About LiveBench Instruction Following

LiveBench tasks that require following detailed formatting and content instructions.

Category
Instruction Following
Introduced
2024
Unit
Percent (random guessing ≈ 0%)
Official site
livebench.ai

Top 15 models

Top models on LiveBench Instruction Following
  1. GPT-5.1 93.3%
  2. Gemini 2.0 Flash (Feb 2025) 85.8%
  3. o3-mini 84.4%
  4. Gemini 2.0 Pro 83.4%
  5. Llama-3.3-70B-Instruct 82.7%
  6. QwQ-32B 81.8%
  7. o1 81.5%
  8. DeepSeek-V3 81.5%
  9. Claude 3.7 Sonnet 81.3%
  10. Gemini 2.5 Pro 80.6%
  11. DeepSeek-R1 80.5%
  12. Gemini 2.0 Flash-Lite 78.3%
  13. Sonar 76.2%
  14. Qwen2.5-Max 75.3%
  15. Gemma 3 27B 74.9%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

LiveBench Instruction Following results by model
#ModelProviderScoreSettingSourceDate
1GPT-5.1 OpenAI93.3%highEpoch AI
2Gemini 2.0 Flash (Feb 2025) Google85.8%Epoch AI
3o3-mini OpenAI84.4%highEpoch AI
4Gemini 2.0 Pro Google83.4%Epoch AI
5Llama-3.3-70B-Instruct Meta82.7%Epoch AI
6QwQ-32B Alibaba (Qwen)81.8%Epoch AI
7o1 OpenAI81.5%highEpoch AI
8DeepSeek-V3 DeepSeek81.5%Epoch AI
9Claude 3.7 Sonnet Anthropic81.3%Epoch AI
10Gemini 2.5 Pro Google80.6%Epoch AI
11DeepSeek-R1 DeepSeek80.5%Epoch AI
12Gemini 2.0 Flash-Lite Google78.3%Epoch AI
13Sonar Perplexity76.2%Epoch AI
14Qwen2.5-Max Alibaba (Qwen)75.3%Epoch AI
15Gemma 3 27B Google74.9%Epoch AI
16GPT-4.5 OpenAI72.3%Epoch AI
17DeepSeek-R1-Distill-Llama-70B DeepSeek69.9%Epoch AI
18Grok-2 (Dec 2024) xAI69.6%Epoch AI
19Claude 3.5 Sonnet Anthropic69.3%Epoch AI
20GPT-4o OpenAI68.6%Epoch AI
21Mistral Large Mistral AI67.9%Epoch AI
22Amazon Nova Pro Amazon67.1%Epoch AI
23o1-mini OpenAI65.4%mediumEpoch AI
24Claude 3 Opus Anthropic63.9%Epoch AI
25Mistral Small Mistral AI63.7%Epoch AI
26Claude 3.5 Haiku Anthropic61.9%Epoch AI
27OLMo 2 Furious 13B Allen Institute for AI (Ai2)60.6%Epoch AI
28Qwen2.5-Coder-32B Alibaba (Qwen)58.7%Epoch AI
29Phi-4 Microsoft58.4%Epoch AI
30Gemma 2 27B Google58.1%Epoch AI
31Command R+ Cohere57.6%Epoch AI
32GPT-4o mini OpenAI56.8%Epoch AI
33DeepSeek-R1-Distill-Qwen-32B DeepSeek55.7%Epoch AI
34Command R Cohere55.6%Epoch AI
35Amazon Nova Lite Amazon54.1%Epoch AI
36Gemma 2 9B Google52.6%Epoch AI
37Amazon Nova Micro Amazon48%Epoch AI
38Phi 3 Small 8k Instruct Microsoft47.2%Epoch AI
39Phi 3 Mini 4k Instruct Microsoft39.1%Epoch AI

Compare the leaders

Other instruction following benchmarks

Frequently asked questions

What does LiveBench Instruction Following measure?

LiveBench tasks that require following detailed formatting and content instructions.

Which model has the highest LiveBench Instruction Following score?

As of October 2026, GPT-5.1 has the highest published LiveBench Instruction Following score on Noometry at 93.3%, out of 39 models with results.

What is the best open-weight model on LiveBench Instruction Following?

Llama-3.3-70B-Instruct has the highest LiveBench Instruction Following accuracy among open-weight models at 82.7%, ranking 5 of 39 overall.