How We Score

Our composite ranking system combines multiple benchmarks into a single, comparable score. Here's how it works.

Composite Score

The composite score is a weighted average of all tracked benchmark results for each model. It provides a single number that captures overall model capability across reasoning, coding, knowledge, and instruction following.

Benchmark Categories

We group benchmarks into four categories, each with its own weight in the composite:

  • Reasoning (35%) — MMLU, GPQA, AIME, MATH. Tests knowledge, mathematical reasoning, and scientific understanding.
  • Coding (30%) — HumanEval, SWE-Bench, LiveCodeBench. Tests code generation, debugging, and real-world software engineering.
  • Chat (20%) — Arena ELO, IFEval. Tests instruction following and human preference alignment.
  • Multimodal (15%) — MMMU. Tests visual reasoning and cross-modal understanding.

Normalization

Raw scores are normalized to a 0–100 scale before weighting. ELO scores are normalized relative to the range of observed values (typically 1100–1400). Percentage-based benchmarks are used as-is.

Performance Metrics

Beyond quality benchmarks, we track operational metrics that matter for production use:

  • Speed — Output tokens per second, measured on standard hardware.
  • Latency — Time to first token (TTFT) and P50/P99 response times.
  • Pricing — Input and output cost per million tokens from official API pricing.

Data Sources

All benchmark data comes from publicly available sources:

  • LMSYS Chatbot Arena for ELO ratings
  • Papers With Code for academic benchmarks
  • Official model cards and technical reports
  • Provider API documentation for pricing
  • Independent evaluations and community testing

Update Frequency

We update our data weekly. Major model releases trigger an immediate update. Each data point on the leaderboard shows its last update date.

Limitations

No benchmark is perfect. Our composite score is an approximation of model capability. We recommend using the detailed per-benchmark breakdowns for specific use cases, and the Compare tool for head-to-head evaluation.