Benchmark

IFEval

Instruction Following Evaluation

Strict instruction-following compliance across formatting and constraint tasks.

25
Models Tested
Claude Opus 4
Top Model
94.2%
Top Score
100 %
Max Possible
#ModelIFEval ScoreProviderComposite
1
Claude Opus 4
Anthropic
94.2%Anthropic97.8
2
o3
OpenAI
92.8%OpenAI97.1
392.3%Google97.5
490.1%Anthropic90.3
5
o4-mini
OpenAI
89.5%OpenAI91.5
6
DeepSeek R1
DeepSeek
88.3%DeepSeek92.3
787.8%xAI88.8
887.4%Meta87.5
9
DeepSeek V3
DeepSeek
86.7%DeepSeek86.1
1086.5%Anthropic88.4
11
GPT-4o
OpenAI
86.1%OpenAI89.2
1285.7%Google86.7
1385.2%Alibaba84.5
1484.3%Meta82.1
1583.5%Meta79.8
1683.2%Google80.5
1782.4%OpenAI75.8
1882.1%Anthropic79.2
19
Mistral Large
Mistral AI
81.2%Mistral AI76.5
2080.1%Alibaba76
2179.8%Meta74.3
2278.5%xAI74.8
23
Mixtral 8x22B
Mistral AI
75.3%Mistral AI65.2
2472.5%Cohere60.8
25
Phi-4
Microsoft
70.2%Microsoft64.5
Source: Zhou et al., 2023