AI 뉴스
모델 출시, 벤치마크 업데이트, 분석을 한 피드로.
연간 데이터 스냅샷: Podium Score 선두는 누구인지, 오픈 웨이트 격차는 얼마나 되는지, 가격은 어디에 안착했는지, 그리고 어떤 벤치마크가 한 해를 결정했는지.
2026년 어떤 AI 모델이 가장 좋은 코드를 작성할까? Claude Fable 5, Claude Mythos Preview, GPT-5.6 Sol, Kimi K3를 SWE-Bench Verified, LiveCodeBench, Terminal-Bench에서 비교합니다.
Kimi K3가 Podium Score 83.3 기록 — 프런티어 독점 모델에 그 어느 때보다 근접. 집계 랭킹 데이터로 2026년 오픈 웨이트 격차를 정량화합니다.
178배의 격차: 2026년 가장 저렴하고 가장 비싼 프런티어 LLM, 지능당 가격 분석, 그리고 가치 스위트스팟의 위치.
Anthropic's next-generation model family, top-ranked on BenchLM with highest composite score.
Gemini 3.5 Flash-Lite가 초당 389 토큰으로 선두. 프런티어 모델의 완전한 속도 랭킹과 속도가 가장 저평가된 벤치마크인 이유.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Compact Inkling variant for fast, cheap inference.
Research preview of Anthropic’s next-generation model family.
Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.
Updated LiveCodeBench scores for all models with latest contamination-free results.
Alibaba’s newest hosted flagship, a fast riser on human preference arenas.
Anthropic frontier model with adaptive reasoning effort levels, leading agentic coding benchmarks.
Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.
Lowest-cost Gemini 3.5 tier for extremely high-throughput serving.
Newest Flash-class Gemini, near-pro intelligence at 200+ tokens/s.
Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.
Moonshot’s frontier MoE with 1M context, top-tier agentic benchmark results.
Thinking Machines’ first frontier model.
실제 사용자의 선호도를 반영하는 라이브 대결 방식 평가 시스템 분석.
LLMPodium now available in 10 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR.
Meta’s proprietary frontier model line, top-10 on blind human preference.
Fast, cheap GPT-5.6 variant built for high-volume production traffic.
Mid-tier GPT-5.6 model balancing intelligence with very high throughput.
Flagship of the GPT-5.6 series for complex reasoning, coding and multi-step agentic workflows.
xAI flagship with real-time knowledge and strong agentic results.
Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.
Tencent’s latest Hunyuan generation model.
데이터 오염을 방지하고 신뢰할 수 있는 평가 체계를 구축하는 방법.
Updated pricing data for all models from official API documentation.
Balanced Claude tier with adaptive reasoning, strong speed-to-intelligence ratio.
Updated SWE-Bench Verified scores for frontier models.
Added DeepSeek R1 reasoning model with chain-of-thought capabilities.
LLMPodium이 여러 공개 리더보드의 점수를 가중치 계산하는 방법.
New Arena page for side-by-side model comparison launched.
Z.ai flagship with strong agent tool use and 150+ tokens/s.
Code-specialized Kimi variant for repository-level engineering.
Anthropic flagship with adaptive reasoning and Opus-class fallback, tuned for the hardest open-ended and agentic tasks.
NVIDIA’s 550B open MoE built for enterprise reasoning pipelines.
Nex AGI’s efficiency-focused frontier model.