Report

State of LLMs 2026: The Podium Report

A snapshot of the model landscape as of August 2026, drawn entirely from the data behind our leaderboard — 58 tracked models, 25 benchmarks, five source leaderboards.

The top of the podium

Claude Mythos Preview holds #1 at 97.4, followed by Claude Fable 5 (93.4) and — the only non-Anthropic model on the podium — Kimi K3 (83.3), which is also the best open-weights model we track. Anthropic occupies four of the top five slots.

The open-weights gap: 14 points and narrowing

Kimi K3's 83.3 sits about 14 Podium points below the proprietary frontier. On LiveCodeBench it beats several closed models that cost ten times more. After the top three open models, however, scores fall below 50 — the open field is currently one lab deep at the frontier.

Pricing: a 178× spread

Output pricing spans $0.28/M (DeepSeek V4 Flash) to $50/M (Claude Fable 5). The market has segmented into three tiers: utility (<$1/M), workhorse ($1–10/M) and frontier ($25–50/M). The value leaderboard shows the intelligence-per-dollar sweet spot sits in the workhorse tier.

Speed: the flash wars

Google's Flash family owns the speed podium — Gemini 3.5 Flash-Lite at 389 tokens/second — while Meta's Muse Spark 1.1 (208 t/s at $4.25/M) is the year's best speed-per-dollar story.

What decided 2026

Agentic benchmarks replaced chat preference as the battleground: SWE-Bench Verified (Fable 5 at 95%) and Terminal-Bench separated the leaders, while Humanity's Last Exam remains the frontier's ceiling — nobody cracks it comfortably. Explore any matchup on the new head-to-head pages.

← All Articles