AI Model Speed Rankings
Models ranked by measured output throughput (tokens per second) and average latency.
- ByteDance Seedance 2.0 Mini — Throughput: 1234.0 tps, Latency: 175775.50s
- ByteDance Doubao-Seedream-4.5 — Throughput: 1115.2 tps, Latency: 12912.00s
- ByteDance Seedance 2.5 — Throughput: 837.3 tps, Latency: 233242.00s
- ByteDance Seedance 2.0 Fast — Throughput: 823.4 tps, Latency: 132255.00s
- ByteDance Doubao-Seedream-4.5 — Throughput: 480.4 tps, Latency: 36188.00s
- Alibaba qwen-mt-turbo — Throughput: 442.4 tps, Latency: 296.00s
- OpenAI GPT-5 — Throughput: 330.7 tps, Latency: 14521.96s
- xAI Grok 4.20 Multi-Agent — Throughput: 293.9 tps, Latency: 0.00s
- Google Gemini 3.6 Flash — Throughput: 278.2 tps, Latency: 5711.86s
- DeepSeek DeepSeek V4.1 Flash — Throughput: 277.8 tps, Latency: 2635.50s
- Google Gemini 3.1 Flash-Lite — Throughput: 274.6 tps, Latency: 3140.04s
- ByteDance Doubao-Seedream-5.0-lite — Throughput: 236.0 tps, Latency: 61021.00s
- ByteDance Doubao-Seedream-5.0-pro — Throughput: 205.3 tps, Latency: 75736.00s
- OpenAI gpt-5-nano — Throughput: 195.5 tps, Latency: 822.59s
- OpenAI o3 — Throughput: 184.6 tps, Latency: 0.00s
- Google Gemini 3.1 Pro Preview — Throughput: 179.1 tps, Latency: 22592.51s
- Google Gemini 3.5 Flash — Throughput: 168.2 tps, Latency: 18623.14s
- Google Gemini 3 Flash Preview — Throughput: 160.1 tps, Latency: 11859.89s
- xAI Grok 4.3 — Throughput: 158.4 tps, Latency: 12835.00s
- OpenAI GPT Image 1.5 — Throughput: 155.2 tps, Latency: 43711.00s
- Google gemini-2.5-flash — Throughput: 151.9 tps, Latency: 35168.00s
- OpenAI dall-e-3 — Throughput: 142.7 tps, Latency: 29253.50s
- xAI grok-3 — Throughput: 133.7 tps, Latency: 616.00s
- Anthropic Claude Claude Haiku 4.5 — Throughput: 130.6 tps, Latency: 4621.16s
- OpenAI gpt-4-gizmo-* — Throughput: 118.8 tps, Latency: 5586.06s
- Anthropic Claude Claude Sonnet 5 — Throughput: 113.9 tps, Latency: 7087.13s
- OpenAI GPT Image 2.5 Dev — Throughput: 109.0 tps, Latency: 44028.70s
- Z.ai GLM 5.2 — Throughput: 105.7 tps, Latency: 628.00s
- OpenAI gpt-5-mini — Throughput: 101.9 tps, Latency: 3849.00s
- Anthropic Claude Claude Opus 4.7 — Throughput: 98.4 tps, Latency: 3037.00s
- Z.ai GLM 5.1 — Throughput: 92.7 tps, Latency: 942.00s
- Google Gemini 3.1 Pro Preview Customtools — Throughput: 91.2 tps, Latency: 42673.33s
- Z.ai GLM 5.3 Flash — Throughput: 88.4 tps, Latency: 416.50s
- Google Gemini 3.7 Flash — Throughput: 88.0 tps, Latency: 11196.83s
- xAI Grok 4.20 — Throughput: 87.4 tps, Latency: 0.00s
- DeepSeek DeepSeek-V4-Flash — Throughput: 87.1 tps, Latency: 4568.90s
- Anthropic Claude Claude Opus 5 — Throughput: 86.8 tps, Latency: 3184.47s
- OpenAI GPT Image 2 Develop — Throughput: 81.7 tps, Latency: 25212.97s
- Google gemini-2.5-pro — Throughput: 81.2 tps, Latency: 50125.00s
- Alibaba qwen3-vl-plus — Throughput: 79.5 tps, Latency: 1297.00s
- xAI Grok 4.5 — Throughput: 78.7 tps, Latency: 11214.97s
- Google Gemini 3.8 Flash — Throughput: 78.4 tps, Latency: 12756.54s
- xAI grok-4-1-fast-reasoning — Throughput: 78.2 tps, Latency: 0.00s
- OpenAI gpt-image-1-mini — Throughput: 76.5 tps, Latency: 13478.00s
- DeepSeek DeepSeek-V4-Pro-0813 — Throughput: 76.0 tps, Latency: 1437.25s
- Z.ai GLM 5.3 — Throughput: 74.4 tps, Latency: 0.00s
- OpenAI Chat Latest (GPT-5.5 Instant) — Throughput: 72.8 tps, Latency: 0.00s
- Anthropic Claude Claude Opus 4.8 — Throughput: 70.1 tps, Latency: 3819.06s
- Thinking Machines Thinking Machines: Inkling — Throughput: 68.6 tps, Latency: 0.00s
- xAI Grok Build 0.1 — Throughput: 67.4 tps, Latency: 0.00s
- Anthropic Claude Claude Opus 4.6 — Throughput: 66.2 tps, Latency: 7395.29s
- DeepSeek DeepSeek-V4-Flash 0731 — Throughput: 65.2 tps, Latency: 1073.82s
- Google Gemini 3.1 Flash-Lite Preview — Throughput: 64.5 tps, Latency: 1027.67s
- Alibaba qwen3-235b-a22b-thinking-2507 — Throughput: 61.5 tps, Latency: 347.00s
- OpenAI GPT Image 2.5 Flare — Throughput: 60.9 tps, Latency: 23470.21s
- Moonshot Kimi K2.6 — Throughput: 60.2 tps, Latency: 964.00s
- Anthropic Claude Claude Sonnet 4.6 — Throughput: 59.2 tps, Latency: 5684.37s
- Z.ai GLM 5 — Throughput: 57.9 tps, Latency: 769.50s
- Moonshot kimi-k2-thinking — Throughput: 56.0 tps, Latency: 419.50s
- xAI Grok 4.6 — Throughput: 55.4 tps, Latency: 12215.67s
- MiniMax MiniMax-M2.7 Highspeed — Throughput: 54.8 tps, Latency: 0.00s
- OpenAI GPT-6 Astra — Throughput: 53.9 tps, Latency: 7807.80s
- DeepSeek DeepSeek-V3.2 — Throughput: 53.2 tps, Latency: 1238.86s
- Anthropic Claude Claude Fable 5 — Throughput: 53.0 tps, Latency: 14680.30s
- OpenAI gpt-3.5-turbo — Throughput: 52.8 tps, Latency: 0.00s
- Z.ai GLM 4.7 — Throughput: 50.3 tps, Latency: 1082.00s
- OpenAI GPT Image 2.5 Sunburst — Throughput: 47.7 tps, Latency: 31516.85s
- Anthropic Claude Claude Fable 5.1 — Throughput: 47.6 tps, Latency: 5663.91s
- OpenAI GPT-5.5 Pro — Throughput: 46.6 tps, Latency: 465.00s
- Anthropic Claude Claude Sonnet 4.5 — Throughput: 46.0 tps, Latency: 1833.95s
- MiniMax MiniMax M3 — Throughput: 46.0 tps, Latency: 1363.70s
- DeepSeek deepseek-v3.1 — Throughput: 45.2 tps, Latency: 2699.00s
- OpenAI gpt-4.1 — Throughput: 44.6 tps, Latency: 18569.09s
- MiniMax MiniMax-M2.1-lightning — Throughput: 44.4 tps, Latency: 0.00s
- OpenAI GPT-5.1 — Throughput: 43.6 tps, Latency: 14343.35s
- xAI grok-3-mini — Throughput: 43.1 tps, Latency: 0.00s
- xAI grok-4-fast-reasoning — Throughput: 42.5 tps, Latency: 0.00s
- Alibaba qwen-plus-latest — Throughput: 42.3 tps, Latency: 0.00s
- xAI grok-4-0709 — Throughput: 42.0 tps, Latency: 4728.44s
- OpenAI GPT-5.6 Luna — Throughput: 40.1 tps, Latency: 5637.80s
- Alibaba qvq-plus — Throughput: 40.0 tps, Latency: 487.00s
- xAI Grok 4.20 (Non-Reasoning) — Throughput: 38.6 tps, Latency: 1513.09s
- Google Nano Banana (Gemini 2.5 Flash Image 🍌) — Throughput: 38.2 tps, Latency: 6913.71s
- DeepSeek deepseek-v3 — Throughput: 37.7 tps, Latency: 2384.42s
- Google Gemini 3.5 Flash-Lite — Throughput: 37.5 tps, Latency: 4774.04s
- ByteDance Doubao-Seed-2.0-Code — Throughput: 36.2 tps, Latency: 0.00s
- OpenAI GPT Image 2 — Throughput: 35.4 tps, Latency: 34376.67s
- OpenAI GPT-5.6 Terra — Throughput: 34.5 tps, Latency: 58728.28s
- Google Gemini 3 Pro Preview — Throughput: 34.1 tps, Latency: 0.00s
- Moonshot Kimi K3 — Throughput: 33.8 tps, Latency: 4962.00s
- Google Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image 🍌) — Throughput: 33.4 tps, Latency: 44708.07s
- Xiaomi MiMo-V2.5 — Throughput: 31.6 tps, Latency: 17348.58s
- OpenAI GPT-5.6 Sol — Throughput: 31.5 tps, Latency: 4496.53s
- OpenAI GPT-5.2 — Throughput: 31.5 tps, Latency: 36212.29s
- OpenAI GPT-5.4 — Throughput: 30.3 tps, Latency: 12145.89s
- Alibaba qwen3-vl-flash — Throughput: 29.4 tps, Latency: 0.00s
- OpenAI GPT Image 1.0 — Throughput: 28.4 tps, Latency: 30244.50s
- OpenAI GPT-5.5 — Throughput: 28.4 tps, Latency: 1561.78s
- OpenAI gpt-4o — Throughput: 28.1 tps, Latency: 5627.43s
- ByteDance Doubao-Seed-2.0-mini — Throughput: 27.5 tps, Latency: 0.00s