AIニュース
モデルリリース、ベンチマーク更新、分析を1つのフィードに。
年次データスナップショット:Podium Scoreのトップは誰か、オープンウェイトの差はどれほどか、価格の着地点、そして1年を決めたベンチマークは何か。
2026年に最も優れたコードを書くAIモデルは?Claude Fable 5、Claude Mythos Preview、GPT-5.6 Sol、Kimi K3をSWE-Bench Verified・LiveCodeBench・Terminal-Benchで比較します。
Kimi K3はPodium Score 83.3を記録 — フロンティアのプロプライエタリモデルに過去最も接近。ランキング集計データで2026年のオープンウェイト格差を定量化します。
178倍の開き:2026年の最安・最高額フロンティアLLM、知能あたりの価格分析、そしてバリューの狙い目。
Anthropic's next-generation model family, top-ranked on BenchLM with highest composite score.
Gemini 3.5 Flash-Liteが389トークン/秒でリード。フロンティアモデルの完全な速度ランキングと、なぜ速度が最も過小評価されたベンチマークなのか。
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Compact Inkling variant for fast, cheap inference.
Research preview of Anthropic’s next-generation model family.
Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.
Updated LiveCodeBench scores for all models with latest contamination-free results.
Alibaba’s newest hosted flagship, a fast riser on human preference arenas.
Anthropic frontier model with adaptive reasoning effort levels, leading agentic coding benchmarks.
Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.
Lowest-cost Gemini 3.5 tier for extremely high-throughput serving.
Newest Flash-class Gemini, near-pro intelligence at 200+ tokens/s.
Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.
Moonshot’s frontier MoE with 1M context, top-tier agentic benchmark results.
Thinking Machines’ first frontier model.
リアルタイムのクラウドソース評価により、人間が実際に好むAIモデルを明らかにする仕組み。
LLMPodium now available in 10 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR.
Meta’s proprietary frontier model line, top-10 on blind human preference.
Fast, cheap GPT-5.6 variant built for high-volume production traffic.
Mid-tier GPT-5.6 model balancing intelligence with very high throughput.
Flagship of the GPT-5.6 series for complex reasoning, coding and multi-step agentic workflows.
xAI flagship with real-time knowledge and strong agentic results.
Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.
Tencent’s latest Hunyuan generation model.
データ汚染を防ぎ、信頼性の高い評価システムを構築する方法。
Updated pricing data for all models from official API documentation.
Balanced Claude tier with adaptive reasoning, strong speed-to-intelligence ratio.
Updated SWE-Bench Verified scores for frontier models.
Added DeepSeek R1 reasoning model with chain-of-thought capabilities.
LLMPodiumが複数の公開リーダーボードのスコアを正規化・加重計算する仕組み。
New Arena page for side-by-side model comparison launched.
Z.ai flagship with strong agent tool use and 150+ tokens/s.
Code-specialized Kimi variant for repository-level engineering.
Anthropic flagship with adaptive reasoning and Opus-class fallback, tuned for the hardest open-ended and agentic tasks.
NVIDIA’s 550B open MoE built for enterprise reasoning pipelines.
Nex AGI’s efficiency-focused frontier model.