AI 新闻
模型发布、基准更新与分析——来自我们的数据管道。
我们的年度数据快照:谁领跑 Podium Score、开放权重差距有多大、价格落在哪里,以及哪些基准测试决定了这一年。
2026年哪些AI模型的代码写得最好?我们对比 Claude Fable 5、Claude Mythos Preview、GPT-5.6 Sol 与 Kimi K3 在 SWE-Bench Verified、LiveCodeBench 和 Terminal-Bench 上的表现。
Kimi K3 的 Podium Score 达到 83.3——比以往任何时候都更接近前沿专有模型。我们用排行榜汇总数据量化2026年的开放权重差距。
178倍的价差:2026年最便宜与最贵的前沿LLM、每点智能的价格分析,以及价值甜区在哪里。
Anthropic's next-generation model family, top-ranked on BenchLM with highest composite score.
Gemini 3.5 Flash-Lite 以 389 token/秒领跑。前沿模型的完整速度排名,以及为什么速度是最被低估的基准。
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Compact Inkling variant for fast, cheap inference.
Research preview of Anthropic’s next-generation model family.
Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.
Updated LiveCodeBench scores for all models with latest contamination-free results.
Alibaba’s newest hosted flagship, a fast riser on human preference arenas.
Anthropic frontier model with adaptive reasoning effort levels, leading agentic coding benchmarks.
Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.
Lowest-cost Gemini 3.5 tier for extremely high-throughput serving.
Newest Flash-class Gemini, near-pro intelligence at 200+ tokens/s.
Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.
Moonshot’s frontier MoE with 1M context, top-tier agentic benchmark results.
Thinking Machines’ first frontier model.
深入了解LLM竞技场概念——通过实时众包评估揭示人们真正偏好的AI模型。
LLMPodium now available in 10 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR.
Meta’s proprietary frontier model line, top-10 on blind human preference.
Fast, cheap GPT-5.6 variant built for high-volume production traffic.
Mid-tier GPT-5.6 model balancing intelligence with very high throughput.
Flagship of the GPT-5.6 series for complex reasoning, coding and multi-step agentic workflows.
xAI flagship with real-time knowledge and strong agentic results.
Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.
Tencent’s latest Hunyuan generation model.
如何防止数据污染、应对提示词敏感性并构建可靠的AI模型评估。
Updated pricing data for all models from official API documentation.
Balanced Claude tier with adaptive reasoning, strong speed-to-intelligence ratio.
Updated SWE-Bench Verified scores for frontier models.
Added DeepSeek R1 reasoning model with chain-of-thought capabilities.
深入分析LLMPodium如何跨多个公开排行榜归一化并加权计算综合得分。
New Arena page for side-by-side model comparison launched.
Z.ai flagship with strong agent tool use and 150+ tokens/s.
Code-specialized Kimi variant for repository-level engineering.
Anthropic flagship with adaptive reasoning and Opus-class fallback, tuned for the hardest open-ended and agentic tasks.
NVIDIA’s 550B open MoE built for enterprise reasoning pipelines.
Nex AGI’s efficiency-focused frontier model.