Benchmark

HumanEval

HumanEval Code Generation

Python function synthesis from docstrings — measures pass@1 accuracy.

25
Models Tested
Claude Opus 4
Top Model
94.5%
Top Score
100 %
Max Possible
#ModelHumanEval ScoreProviderComposite
1
Claude Opus 4
Anthropic
94.5%Anthropic97.8
293.7%Anthropic90.3
393%Google97.5
4
o3
OpenAI
92.7%OpenAI97.1
5
DeepSeek R1
DeepSeek
92.1%DeepSeek92.3
692%Anthropic88.4
7
o4-mini
OpenAI
91.5%OpenAI91.5
890.8%Meta87.5
990.5%xAI88.8
10
GPT-4o
OpenAI
90.2%OpenAI89.2
1189.8%Google86.7
12
DeepSeek V3
DeepSeek
89.4%DeepSeek86.1
1389%Meta82.1
1488.7%Alibaba84.5
1588.1%Anthropic79.2
1687.6%Google80.5
1787.2%OpenAI75.8
1886.4%Meta79.8
19
Mistral Large
Mistral AI
84.3%Mistral AI76.5
2082.4%Alibaba76
2182.1%xAI74.8
2280.5%Meta74.3
23
Mixtral 8x22B
Mistral AI
75.8%Mistral AI65.2
24
Phi-4
Microsoft
72.4%Microsoft64.5
2568.2%Cohere60.8
Source: Chen et al., 2021