Rankings

The Fastest LLMs in 2026: Output Speed and Latency Compared

Speed is the benchmark users feel every day, and in 2026 the spread is enormous: from 389 tokens/second at the top to single digits for large reasoning models. Here is the current state of play from our speed leaderboard.

Fastest models right now

  1. Gemini 3.5 Flash-Lite — 389 t/s
  2. Gemini 3.5 Flash — 267 t/s
  3. Gemini 3.6 Flash — 233 t/s
  4. Muse Spark 1.1 (Meta) — 208 t/s
  5. Qwen3.7 Max — 203 t/s

Google's Flash family occupies the entire podium, and note Meta's Muse Spark 1.1: 208 t/s at $4.25/M output makes it one of the best speed-per-dollar options we track.

Speed vs intelligence trade-off

The fastest models are not the smartest: Flash-Lite trades raw capability for throughput. The practical pattern in 2026 is routing — a fast model for drafts and tool calls, a frontier model like Claude Mythos Preview for the hard 5%. Our comparison tool shows speed and Podium Score side by side.

Latency (TTFT) matters for chat

Perceived responsiveness is dominated by time-to-first-token. For chat products, a 200 t/s model with 150 ms TTFT feels faster than a 389 t/s model with a 2 s cold start — always evaluate both numbers, visible in each model's profile (e.g. Claude Fable 5: 71 t/s, 141 ms TTFT).

Bottom line

For pure throughput pick Gemini Flash-Lite; for balanced speed + quality pick Muse Spark 1.1 or Qwen3.7 Max. Live numbers on the speed leaderboard.

← All Articles