بینچ مارک کا جائزہ
673 ماڈلز پر تمام 25 بینچ مارکس کی تعریفیں اور موجودہ سرکردہ۔
GPQA Diamond
🧠 ReasoningGraduate-level science questions in physics, chemistry and biology, designed to be Google-proof.
Humanity’s Last Exam
🧠 ReasoningThousands of expert-written questions across dozens of fields, the hardest closed-ended academic benchmark in wide use.
ARC-AGI-2
🧠 ReasoningSecond-generation abstraction and reasoning tasks measuring fluid intelligence on novel visual patterns.
AIME 2025
📐 MathAmerican Invitational Mathematics Examination problems solved without tool assistance.
FrontierMath
📐 MathResearch-grade mathematics problems crafted by professional mathematicians.
MATH
📐 MathCompetition mathematics dataset spanning algebra to number theory.
SWE-Bench Verified
💻 CodingReal GitHub issues resolved end-to-end; the industry standard for agentic coding.
SWE-Bench Pro
💻 CodingHarder, contamination-resistant successor to SWE-Bench with commercial-repo issues.
LiveCodeBench
💻 CodingContinuously refreshed competitive programming problems immune to data contamination.
SciCode
💻 CodingScientific coding problems requiring domain knowledge and multi-function solutions.
Terminal-Bench Hard
💻 CodingMulti-step command-line tasks in realistic terminal environments.
HumanEval
💻 CodingFunction-level Python completion, the classic code-generation benchmark.
OSWorld
🤖 AgenticReal computer-use tasks across operating systems: browsers, office apps and file management.
Toolathlon
🤖 AgenticMulti-tool orchestration tasks across real-world APIs.
MCP Atlas
🤖 AgenticTool calling over the Model Context Protocol across complex server graphs.
τ²-Bench Retail
🤖 AgenticDual-control conversational agent tasks in a simulated retail environment.
Apex Agents
🤖 AgenticLong-horizon professional agent tasks with multi-app workflows.
SimpleQA
📚 KnowledgeShort fact-seeking questions measuring hallucination rates.
MMLU-Pro
📚 KnowledgeHarder 14-subject version of MMLU with ten-way multiple choice.
Multilingual MMLU
📚 KnowledgeMMLU translated across 14 languages.
MMMU
👁️ MultimodalMassive Multi-discipline Multimodal Understanding with college-level image reasoning.
MMMU-Pro
👁️ MultimodalHarder multimodal reasoning successor of MMMU.
CharXiv Reasoning
👁️ MultimodalReasoning over scientific charts and figures from arXiv papers.
ScreenSpot-Pro
👁️ MultimodalGUI grounding on high-resolution professional application screens.
MRCR v2
📜 Long ContextMulti-round coreference resolution over needle-in-haystack long contexts.