Speed is the benchmark users feel every day, and in 2026 the spread is enormous: from 389 tokens/second at the top to single digits for large reasoning models. Here is the current state of play from our speed leaderboard.
Fastest models right now
- Gemini 3.5 Flash-Lite — 389 t/s
- Gemini 3.5 Flash — 267 t/s
- Gemini 3.6 Flash — 233 t/s
- Muse Spark 1.1 (Meta) — 208 t/s
- Qwen3.7 Max — 203 t/s
Google's Flash family occupies the entire podium, and note Meta's Muse Spark 1.1: 208 t/s at $4.25/M output makes it one of the best speed-per-dollar options we track.
Speed vs intelligence trade-off
The fastest models are not the smartest: Flash-Lite trades raw capability for throughput. The practical pattern in 2026 is routing — a fast model for drafts and tool calls, a frontier model like Claude Mythos Preview for the hard 5%. Our comparison tool shows speed and Podium Score side by side.
Latency (TTFT) matters for chat
Perceived responsiveness is dominated by time-to-first-token. For chat products, a 200 t/s model with 150 ms TTFT feels faster than a 389 t/s model with a 2 s cold start — always evaluate both numbers, visible in each model's profile (e.g. Claude Fable 5: 71 t/s, 141 ms TTFT).
Bottom line
For pure throughput pick Gemini Flash-Lite; for balanced speed + quality pick Muse Spark 1.1 or Qwen3.7 Max. Live numbers on the speed leaderboard.