👨‍💻 Best AI for Programming

Compare AI models for software development — code generation, debugging, code review, refactoring, and architecture design.Rankings combine multiple benchmark scores weighted by relevance to this use case. Updated August 2026.

Why This Matters

Programming is one of the most common AI workloads — and the best model depends heavily on the specific task. A model that excels at creative writing may struggle with structured data extraction, and vice versa.

We built this use-case ranking by combining multiple benchmark categories with weights tuned to match real-world usage patterns. For programming, the score emphasizes the benchmarks that matter most: task accuracy, output quality, and consistency. Speed and cost are factored in but secondary to quality.

All scores are from public independent benchmarks. For a broader view across all tasks, see the overall LLM leaderboard or the expert picks page.

Quick Answer

The best AI models for programming are:1. Claude Fable 5 (Anthropic, score: 85.6), 2. Kimi K3 (Moonshot AI, score: 76.7), 3. Claude Mythos Preview (Anthropic, score: 70.1).

#ModelProviderScoreSpeedPrice (output)
1Claude Fable 5Anthropic85.671 t/s$50/M
2Kimi K3Moonshot AI76.737 t/s$15/M
3Claude Mythos PreviewAnthropic70.180 t/s$15/M
4Claude Opus 4.8Anthropic68.850 t/s$15/M
5Claude Mythos 5Anthropic66.750 t/s$50/M
6GPT-5.6 SolOpenAI65.572 t/s$30/M
7GPT-5.5OpenAI63.360 t/s$10/M
8Claude Opus 4.7Anthropic61.580 t/s$15/M
9Claude Opus 5Anthropic59.960 t/s$25/M
10Gemini 3.5 FlashGoogle59.9267 t/s$9/M
11Claude Sonnet 4.6 (max)Anthropic57.550 t/s$15/M
12GLM-5.1Z.ai57.374 t/s$4.4/M
13Gemini 3.1 ProGoogle56.2136 t/s$12/M
14MiniMax-M2.7MiniMax55.540 t/s$2/M
15GPT-5.6 TerraOpenAI54.5127 t/s$12/M