Skip to main content
LLM
Podium
排行榜
🏆 Overall Leaderboard
Composite index across all flagship LLMs
💻 Coding & SWE-Bench
Real-world programming & software agent benchmark
🧠 Deep Reasoning
GPQA Diamond, AIME & complex logic
🤖 Autonomous Agents
OSWorld, Toolathlon & multi-turn execution
🔓 Open Weights
DeepSeek, Qwen, Llama & Mistral
竞技场
⚔️ Human Arena Battles
Blind pairwise Elo rankings (700+ models)
⚖️ Side-by-Side Compare
Head-to-head metric comparison of up to 4 models
💰 Cost-per-Task Calculator
Estimate API economics across real workflows
🎯 Model Recommender
Find the optimal model for your budget and speed
模型列表
📦 Model Directory
Detailed profiles, context windows & pricing
🏢 AI Labs & Providers
OpenAI, Anthropic, Google, DeepSeek, Meta
📊 Benchmark Matrix
Evaluation methodologies and leader tables
nav.resources
📐 Podium Methodology
Mathematical aggregation formula explained
📰 Changelog & News
Daily model updates and new evaluations
✍️ Research & Articles
In-depth AI benchmarking whitepapers
❓ Frequently Asked Questions
Common questions on scores and ranking
📖 LLM Glossary
Definitions of TTFT, TPS, Elo, MoE & tokens
nav.search_hint
⌘K
🇨🇳
ZH
🇺🇸 English
EN
🇨🇳 中文
ZH
🇮🇳 हिन्दी
HI
🇪🇸 Español
ES
🇫🇷 Français
FR
🇸🇦 العربية
AR
🇧🇩 বাংলা
BN
🇧🇷 Português
PT
🇷🇺 Русский
RU
🇵🇰 اردو
UR
🇮🇩 Bahasa Indonesia
ID
🇩🇪 Deutsch
DE
🇯🇵 日本語
JA
🇰🇷 한국어
KO
🇹🇭 ไทย
TH
🇮🇹 Italiano
IT
排行榜
🏆 排行榜
Overall Composite Leaderboard
Coding & Software Agents
Deep Reasoning & Math
Autonomous Agent Swarms
Open Weights & Self-Hosted
⚔️ 竞技场 & Compare
Human Battle Arena (700+ Models)
Side-by-Side Model Compare
Cost-per-Task Calculator
Interactive Model Finder
📦 Directories & Evals
Model Specifications Directory
AI Labs & Cloud Providers
Benchmark Methodologies
📖 Research & Info
Scoring Methodology
Research Blog & Articles
Changelog & New Models
Frequently Asked Questions
Technical LLM Glossary
About LLMPodium
🌐 Language / Язык / 语言
🇺🇸
English
🇨🇳
中文
🇮🇳
हिन्दी
🇪🇸
Español
🇫🇷
Français
🇸🇦
العربية
🇧🇩
বাংলা
🇧🇷
Português
🇷🇺
Русский
🇵🇰
اردو
🇮🇩
Bahasa Indonesia
🇩🇪
Deutsch
🇯🇵
日本語
🇰🇷
한국어
🇹🇭
ไทย
🇮🇹
Italiano
Explore Full Leaderboard →
首页
/
比较
/
OpenAI vs Anthropic
OpenAI vs Anthropic
OpenAI 与 Anthropic 正面交锋:模型、Podium Score、基准、速度与价格全面对比。 更新于2026年8月。
对比
75
平均分
$14.00
平均价格
5
模型
VS
85
平均分
$24.00
平均价格
5
模型
OpenAI 热门模型
1
GPT-5.6 Sol
80.8
2
GPT 5.6 Sol Xhigh
74.0
3
GPT 5.5 High
74.0
4
GPT 5.4 High
74.0
5
GPT 5.2 Chat Latest 20260210
74.0
Anthropic 热门模型
1
Claude Mythos Preview
97.4
2
Claude Fable 5
93.4
3
Claude Opus 5
81.4
4
Claude Opus 4.8
76.0
5
Claude Opus 4 6 Thinking
75.0
相关比较
OpenAI vs Google
Anthropic vs Google
OpenAI vs DeepSeek
Claude vs Gemini