# LLMPodium > LLMPodium is the definitive open platform and leaderboard for evaluating, comparing, and discovering large language models. > We aggregate five independent evaluation authorities (Arena.ai, Artificial Analysis, LLM Stats, Vellum, LLMBase) > into a normalized, single Podium Score (0–100) for 58 flagship models, track 25+ benchmarks across 11 domains, > and mirror the full human-preference arena table with 719+ models. Data refreshed continuously. ## Key Pages & Endpoints - [Overall LLM Leaderboard](https://llmpodium.com/leaderboard): main rankings by Podium Score (sortable, filterable by provider, license, modality) - [Best LLMs for Coding](https://llmpodium.com/leaderboard/coding): SWE-Bench Pro, LiveCodeBench, Terminal-Bench - [Best LLMs for Math](https://llmpodium.com/leaderboard/math): AIME 2025, FrontierMath, Math500 - [Best Reasoning Models](https://llmpodium.com/leaderboard/reasoning): GPQA Diamond, HLE, ARC-AGI-2 - [Best Agentic Models](https://llmpodium.com/leaderboard/agentic): OSWorld, τ²-Bench, Toolathlon - [Best Knowledge Models](https://llmpodium.com/leaderboard/knowledge): SimpleQA, MMLU-Pro - [Best Multimodal Models](https://llmpodium.com/leaderboard/multimodal): MMMU, CharXiv-R, ScreenSpot-Pro - [Best Long-Context Models](https://llmpodium.com/leaderboard/long-context): MRCR v2, Needle in a Haystack - [Fastest LLMs by Output Speed](https://llmpodium.com/leaderboard/speed): Tokens per second (TPS) & Time to First Token (TTFT) - [Best Price-Performance LLMs](https://llmpodium.com/leaderboard/value): Cost per 1M tokens vs Benchmark Quality - [Best Open-Weight Models](https://llmpodium.com/leaderboard/open): DeepSeek, Qwen, Llama 4, Mistral - [Best Tool-Calling LLMs](https://llmpodium.com/leaderboard/tool-calling): Toolathlon, MCP Atlas, τ²-Bench - [Best Coding Agent LLMs](https://llmpodium.com/leaderboard/coding-agents): SWE-Bench Pro, Terminal-Bench - [Full Arena Table](https://llmpodium.com/arena): 719+ models with Elo scores, confidence intervals, votes, pricing - [AI Models Directory](https://llmpodium.com/models): 58+ flagship model profiles with full parameter specs - [AI Providers Directory](https://llmpodium.com/providers): Lab hubs for OpenAI, Anthropic, Google, DeepSeek, Meta, Mistral, xAI & more - [Side-by-Side Compare](https://llmpodium.com/compare): Compare up to 4 models across 30+ metrics - [Model Head-to-Head Comparisons](https://llmpodium.com/vs/claude-mythos-preview-vs-claude-fable-5): Pairwise comparison pages - [Cost per Task Calculator](https://llmpodium.com/cost-per-task): Economic calculations for coding, extraction, agent swarms - [Model Recommender](https://llmpodium.com/recommender): Interactive quiz to find the optimal model - [Benchmark Matrix & Definitions](https://llmpodium.com/benchmarks): All 25+ benchmark methodologies - [Podium Score Methodology](https://llmpodium.com/methodology): Mathematical aggregation formula - [LLM Glossary](https://llmpodium.com/glossary): Technical dictionary of LLM metrics & concepts - [FAQ](https://llmpodium.com/faq): Common questions on evaluation and rankings - [Research Blog](https://llmpodium.com/blog): Methodology whitepapers, evaluation best practices - [News](https://llmpodium.com/news): Changelog and newly benchmarked models - [RSS Feed](https://llmpodium.com/rss.xml): Automated feed of new models and score updates - [Open JSON API](https://llmpodium.com/api/data/compare.json): Machine-readable evaluations ## Top Models (by Composite Podium Score) 1. Claude Mythos Preview (Anthropic) 2. Claude Fable 5 (Anthropic) 3. Kimi K3 (Moonshot AI) 4. Claude Opus 5 (Anthropic) 5. GPT-5.6 Sol (OpenAI) ## Podium Score Formula Composite Score = 0.35 · Arena Elo + 0.30 · Benchmark Average + 0.20 · Intelligence Index + 0.15 · LLM Stats Score (Min-max normalized across all models; missing values dynamically re-weighted). ## Evaluated Labs & Providers OpenAI, Anthropic, Google DeepMind, Meta AI, xAI, DeepSeek, Mistral AI, Moonshot AI, Alibaba Cloud, Cohere, Microsoft, Amazon Bedrock, Baidu, ByteDance, Zhipu AI, 01.AI, Reka, AI21 Labs. ## Supported Languages (16 Locales) English (en), 中文 (zh), 日本語 (ja), 한국어 (ko), ไทย (th), Русский (ru), Deutsch (de), Español (es), Italiano (it), Français (fr), हिन्दी (hi), العربية (ar), বাংলা (bn), Português (pt), اردو (ur), Bahasa Indonesia (id). ## About LLMPodium is maintained as an independent, transparent AI benchmarking observatory. Website: https://llmpodium.com GitHub: https://github.com/yakushevhk/LLMPodium