Skip to main content
LLMPodium
Leaderboard
🏆 Overall LeaderboardComposite index across all flagship LLMs💻 Coding & SWE-BenchReal-world programming & software agent benchmark🧠 Deep ReasoningGPQA Diamond, AIME & complex logic🤖 Autonomous AgentsOSWorld, Toolathlon & multi-turn execution🔓 Open WeightsDeepSeek, Qwen, Llama & Mistral
Arena
⚔️ Human Arena BattlesBlind pairwise Elo rankings (700+ models)⚖️ Side-by-Side CompareHead-to-head metric comparison of up to 4 models💰 Cost-per-Task CalculatorEstimate API economics across real workflows🎯 Model RecommenderFind the optimal model for your budget and speed
Models
📦 Model DirectoryDetailed profiles, context windows & pricing🏢 AI Labs & ProvidersOpenAI, Anthropic, Google, DeepSeek, Meta📊 Benchmark MatrixEvaluation methodologies and leader tables
nav.resources
📐 Podium MethodologyMathematical aggregation formula explained📰 Changelog & NewsDaily model updates and new evaluations✍️ Research & ArticlesIn-depth AI benchmarking whitepapers❓ Frequently Asked QuestionsCommon questions on scores and ranking📖 LLM GlossaryDefinitions of TTFT, TPS, Elo, MoE & tokens
🇺🇸 EnglishEN🇨🇳 中文ZH🇮🇳 हिन्दीHI🇪🇸 EspañolES🇫🇷 FrançaisFR🇸🇦 العربيةAR🇧🇩 বাংলাBN🇧🇷 PortuguêsPT🇷🇺 РусскийRU🇵🇰 اردوUR🇮🇩 Bahasa IndonesiaID🇩🇪 DeutschDE🇯🇵 日本語JA🇰🇷 한국어KO🇹🇭 ไทยTH🇮🇹 ItalianoIT
Leaderboard
🏆 Leaderboard
Overall Composite LeaderboardCoding & Software AgentsDeep Reasoning & MathAutonomous Agent SwarmsOpen Weights & Self-Hosted
⚔️ Arena & Compare
Human Battle Arena (700+ Models)Side-by-Side Model CompareCost-per-Task CalculatorInteractive Model Finder
📦 Directories & Evals
Model Specifications DirectoryAI Labs & Cloud ProvidersBenchmark Methodologies
📖 Research & Info
Scoring MethodologyResearch Blog & ArticlesChangelog & New ModelsFrequently Asked QuestionsTechnical LLM GlossaryAbout LLMPodium
🌐 Language / Язык / 语言
🇺🇸English🇨🇳中文🇮🇳हिन्दी🇪🇸Español🇫🇷Français🇸🇦العربية🇧🇩বাংলা🇧🇷Português🇷🇺Русский🇵🇰اردو🇮🇩Bahasa Indonesia🇩🇪Deutsch🇯🇵日本語🇰🇷한국어🇹🇭ไทย🇮🇹Italiano
Explore Full Leaderboard →
  1. Home
  2. /AI Providers
  3. /Z.ai

Z.ai

23 models from Z.ai, ranked by Podium Score.

Glm 5.2 Max
Top model
73.0
Podium Score
51.7
Average score
11
open weights
Glm 5.2 Max
73.0
128K ctx50 t/s$3/M out
73.0
Glm 4.7
72.0
128K ctx50 t/s$3/M out
72.0
Glm 5v Turbo
72.0
128K ctx50 t/s$3/M out
72.0
Glm 4.6
71.0
128K ctx50 t/s$3/M out
71.0
Glm 4.5
71.0
128K ctx50 t/s$3/M out
71.0
Glm 4.6v
69.0
128K ctx50 t/s$3/M out
69.0
Glm 4.5 Air
69.0
128K ctx50 t/s$3/M out
69.0
Glm 4.7 Flash
68.0
128K ctx50 t/s$3/M out
68.0
Glm 4.5v
68.0
128K ctx50 t/s$3/M out
68.0
Glm 4 Plus 0111
67.0
128K ctx50 t/s$3/M out
67.0
Glm 4 Plus
66.0
128K ctx50 t/s$3/M out
66.0
Glm 4 0520
64.0
128K ctx50 t/s$3/M out
64.0
GLM-5.2
59.1
262K ctx164 t/s$4.29/M out
59.1
Open
Glm 5
52.6
131K ctx60 t/s$3/M out
52.6
Open
GLM-5.1
46.0
262K ctx74 t/s$4.4/M out
46.0
Open
Glm 5 1 Non Reasoning
36.3
131K ctx40 t/s$3/M out
36.3
Open
Glm 5 2 Non Reasoning
34.8
131K ctx40 t/s$3/M out
34.8
Open
Glm 5 Non Reasoning
33.2
131K ctx40 t/s$3/M out
33.2
Open
Glm 4 6 Reasoning
29.3
131K ctx40 t/s$3/M out
29.3
Open
Glm 4 7 Non Reasoning
27.1
131K ctx40 t/s$3/M out
27.1
Open
Glm 4 6v Reasoning
16.9
131K ctx40 t/s$3/M out
16.9
Open
Glm 4 7 Flash Non Reasoning
15.6
131K ctx40 t/s$3/M out
15.6
Open
Glm 4 5v Reasoning
9.0
131K ctx40 t/s$3/M out
9.0
Open
LLMPodium

The definitive open LLM benchmark aggregator and leaderboard. Continuous evaluations across 719+ models and 25+ benchmarks.

Leaderboards synced daily
Leaderboards
  • Leaderboard
  • 💻 Coding
  • 🧠 Reasoning
  • 📐 Math
  • 🤖 Agentic
  • 🔓 Open Weights
  • ⚡ Speed
  • 💰 Value
Tools
  • Arena (719+ models)
  • Compare Side-by-Side
  • Benchmarks Matrix
  • Finder
  • Models Directory
  • AI Providers
  • Cost per Task Calculator
  • #1 Ranked Model Profile
Resources
  • News Feed
  • Blog & Research
  • Best AI Models 2026
  • Enterprise AI Use Cases
  • Methodology
  • Glossary
  • FAQ
  • About
© 2026 LLMPodium. All rights reserved.·Data is regularly updated from public benchmarks.
llms.txtRSS FeedOpen API