Ikhtisar benchmark

Definisi dan pemimpin terkini untuk semua 25 benchmark pada 673 model.

GPQA Diamond

๐Ÿง  Reasoning

Graduate-level science questions in physics, chemistry and biology, designed to be Google-proof.

๐Ÿฅ‡Claude Mythos Preview โ€” 94.6%๐ŸฅˆGPT-5.6 Sol โ€” 94.6%๐Ÿฅ‰Gemini 3.1 Pro โ€” 94.3%

Humanityโ€™s Last Exam

๐Ÿง  Reasoning

Thousands of expert-written questions across dozens of fields, the hardest closed-ended academic benchmark in wide use.

๐Ÿฅ‡Claude Mythos Preview โ€” 64.7%๐ŸฅˆMuse Spark โ€” 58.4%๐Ÿฅ‰Claude Opus 4.8 โ€” 57.9%

ARC-AGI-2

๐Ÿง  Reasoning

Second-generation abstraction and reasoning tasks measuring fluid intelligence on novel visual patterns.

๐Ÿฅ‡GPT-5.5 โ€” 85%๐ŸฅˆGemini 3.1 Pro โ€” 77.1%๐Ÿฅ‰GPT-5.4 โ€” 73.3%

AIME 2025

๐Ÿ“ Math

American Invitational Mathematics Examination problems solved without tool assistance.

๐Ÿฅ‡Grok 4 Heavy โ€” 100%๐ŸฅˆGemini 3 Pro โ€” 100%๐Ÿฅ‰GPT-5.2 โ€” 100%

FrontierMath

๐Ÿ“ Math

Research-grade mathematics problems crafted by professional mathematicians.

๐Ÿฅ‡GPT-5.6 Sol โ€” 89%๐ŸฅˆGPT-5.6 Terra โ€” 84.9%๐Ÿฅ‰GPT-5.6 Luna โ€” 78.6%

MATH

๐Ÿ“ Math

Competition mathematics dataset spanning algebra to number theory.

๐Ÿฅ‡MiMo V2.5 Pro โ€” 86.2%๐ŸฅˆGPT-5 โ€” 84.7%

SWE-Bench Verified

๐Ÿ’ป Coding

Real GitHub issues resolved end-to-end; the industry standard for agentic coding.

๐Ÿฅ‡Claude Fable 5 โ€” 95%๐ŸฅˆClaude Mythos Preview โ€” 93.9%๐Ÿฅ‰Claude Opus 4.8 โ€” 88.6%

SWE-Bench Pro

๐Ÿ’ป Coding

Harder, contamination-resistant successor to SWE-Bench with commercial-repo issues.

๐Ÿฅ‡Claude Mythos Preview โ€” 77.8%๐ŸฅˆClaude Opus 4.8 โ€” 69.2%๐Ÿฅ‰Qwen3.8 Max โ€” 67.7%

LiveCodeBench

๐Ÿ’ป Coding

Continuously refreshed competitive programming problems immune to data contamination.

๐Ÿฅ‡DeepSeek V4 Pro โ€” 93.5%๐ŸฅˆGemini 3 Pro โ€” 91.7%๐Ÿฅ‰DeepSeek V4 Flash โ€” 91.6%

SciCode

๐Ÿ’ป Coding

Scientific coding problems requiring domain knowledge and multi-function solutions.

๐Ÿฅ‡Claude Fable 5 โ€” 60.2%๐ŸฅˆGemini 3.1 Pro โ€” 59%๐Ÿฅ‰Kimi K3 โ€” 58.7%

Terminal-Bench Hard

๐Ÿ’ป Coding

Multi-step command-line tasks in realistic terminal environments.

๐Ÿฅ‡GPT-5.6 Sol โ€” 65.9%๐ŸฅˆClaude Fable 5 โ€” 62.9%๐Ÿฅ‰GPT-5.5 โ€” 60.6%

HumanEval

๐Ÿ’ป Coding

Function-level Python completion, the classic code-generation benchmark.

๐Ÿฅ‡GPT-5 โ€” 93.4%

OSWorld

๐Ÿค– Agentic

Real computer-use tasks across operating systems: browsers, office apps and file management.

๐Ÿฅ‡Claude Opus 4.6 โ€” 72.7%๐ŸฅˆClaude Sonnet 4.6 โ€” 72.5%

Toolathlon

๐Ÿค– Agentic

Multi-tool orchestration tasks across real-world APIs.

๐Ÿฅ‡Kimi K3 โ€” 73.2%๐ŸฅˆClaude Opus 4.8 โ€” 59.9%๐Ÿฅ‰GPT-5.6 Sol โ€” 58%

MCP Atlas

๐Ÿค– Agentic

Tool calling over the Model Context Protocol across complex server graphs.

๐Ÿฅ‡Kimi K3 โ€” 84.2%๐ŸฅˆClaude Opus 4.8 โ€” 82.2%๐Ÿฅ‰Hunyuan Hy3 โ€” 79.1%

ฯ„ยฒ-Bench Retail

๐Ÿค– Agentic

Dual-control conversational agent tasks in a simulated retail environment.

๐Ÿฅ‡GLM-5.2 โ€” 99.1%๐ŸฅˆClaude Fable 5 โ€” 98.5%๐Ÿฅ‰Qwen3.6 Plus โ€” 97.7%

Apex Agents

๐Ÿค– Agentic

Long-horizon professional agent tasks with multi-app workflows.

๐Ÿฅ‡Kimi K3 โ€” 37.6%๐ŸฅˆGemini 3.1 Pro โ€” 33.5%๐Ÿฅ‰Kimi K2.6 โ€” 27.9%

SimpleQA

๐Ÿ“š Knowledge

Short fact-seeking questions measuring hallucination rates.

๐Ÿฅ‡Gemini 3 Pro โ€” 72.1%๐ŸฅˆGemini 3 Flash โ€” 68.7%๐Ÿฅ‰DeepSeek V4 Pro โ€” 57.9%

MMLU-Pro

๐Ÿ“š Knowledge

Harder 14-subject version of MMLU with ten-way multiple choice.

๐Ÿฅ‡Gemini 3 Pro โ€” 89.8%๐ŸฅˆQwen3.7 Max โ€” 89.6%๐Ÿฅ‰Qwen3.7 Plus โ€” 88.5%

Multilingual MMLU

๐Ÿ“š Knowledge

MMLU translated across 14 languages.

๐Ÿฅ‡Claude Mythos Preview โ€” 92.7%๐ŸฅˆGemini 3.1 Pro โ€” 92.6%๐Ÿฅ‰Gemini 3 Pro โ€” 91.8%

MMMU

๐Ÿ‘๏ธ Multimodal

Massive Multi-discipline Multimodal Understanding with college-level image reasoning.

๐Ÿฅ‡Qwen3.6 Plus โ€” 86%๐ŸฅˆGPT-5.1 โ€” 85.4%๐Ÿฅ‰GPT-5 โ€” 84.2%

MMMU-Pro

๐Ÿ‘๏ธ Multimodal

Harder multimodal reasoning successor of MMMU.

๐Ÿฅ‡GPT-5.5 โ€” 83.2%๐ŸฅˆGPT-5.6 Sol โ€” 83%๐Ÿฅ‰Qwen3.8 Max โ€” 82.3%

CharXiv Reasoning

๐Ÿ‘๏ธ Multimodal

Reasoning over scientific charts and figures from arXiv papers.

๐Ÿฅ‡Claude Mythos Preview โ€” 93.2%๐ŸฅˆKimi K3 โ€” 91.3%๐Ÿฅ‰Claude Opus 4.7 โ€” 91%

ScreenSpot-Pro

๐Ÿ‘๏ธ Multimodal

GUI grounding on high-resolution professional application screens.

๐Ÿฅ‡Claude Opus 4.8 โ€” 87.9%๐ŸฅˆGPT-5.2 โ€” 86.3%๐Ÿฅ‰Qwen3.8 Max โ€” 84.5%

MRCR v2

๐Ÿ“œ Long Context

Multi-round coreference resolution over needle-in-haystack long contexts.

๐Ÿฅ‡GPT-5.6 Sol โ€” 91.5%๐ŸฅˆGPT-5.6 Terra โ€” 89.6%๐Ÿฅ‰Claude Opus 4.6 โ€” 76%