🧠 Reasoning LLMリーダーボード

Mathematical reasoning, science, and logical problem solving.

#モデル総合MMLUGPQAAIMEMATH速度
1
Claude Opus 4
Anthropic
97.892.1%83.5%87.4%96.1%45 t/s
297.592%84%92%97.2%85 t/s
3
o3
OpenAI
97.191.2%81.3%91.6%96.7%58 t/s
4
DeepSeek R1
DeepSeek
92.390.8%71.5%79.8%97.3%32 t/s
5
o4-mini
OpenAI
91.588.3%72.1%79.4%93.2%145 t/s
690.390.4%70.2%40.5%85.6%79 t/s
7
GPT-4o
OpenAI
89.288.7%53.6%9.3%76.6%110 t/s
888.889.2%68.5%55.3%88.1%55 t/s
988.488.7%65%16%78.3%82 t/s
1087.589.5%62.8%45.3%84.2%120 t/s
11
DeepSeek V3
DeepSeek
86.188.5%59.1%39.2%90.2%68 t/s
1284.587.8%58.3%35.6%88.5%95 t/s
1382.187.3%51.1%7.8%73.8%29 t/s
1479.886%48.7%8.2%72.1%65 t/s
15
Mistral Large
Mistral AI
76.584%46.5%5.6%69.2%48 t/s
167685.2%45.6%6.8%71.5%55 t/s
17
Phi-4
Microsoft
64.578.5%38.5%3.8%62.5%125 t/s