AI Model Speed Rankings

Models ranked by measured output throughput (tokens per second) and average latency.

  1. ByteDance Seedance 2.0 Mini — Throughput: 1234.0 tps, Latency: 175775.50s
  2. ByteDance Doubao-Seedream-4.5 — Throughput: 1115.2 tps, Latency: 12912.00s
  3. ByteDance Seedance 2.5 — Throughput: 837.3 tps, Latency: 233242.00s
  4. ByteDance Seedance 2.0 Fast — Throughput: 823.4 tps, Latency: 132255.00s
  5. ByteDance Doubao-Seedream-4.5 — Throughput: 480.4 tps, Latency: 36188.00s
  6. Alibaba qwen-mt-turbo — Throughput: 442.4 tps, Latency: 296.00s
  7. OpenAI GPT-5 — Throughput: 330.7 tps, Latency: 14521.96s
  8. xAI Grok 4.20 Multi-Agent — Throughput: 293.9 tps, Latency: 0.00s
  9. Google Gemini 3.6 Flash — Throughput: 278.2 tps, Latency: 5711.86s
  10. DeepSeek DeepSeek V4.1 Flash — Throughput: 277.8 tps, Latency: 2635.50s
  11. Google Gemini 3.1 Flash-Lite — Throughput: 274.6 tps, Latency: 3140.04s
  12. ByteDance Doubao-Seedream-5.0-lite — Throughput: 236.0 tps, Latency: 61021.00s
  13. ByteDance Doubao-Seedream-5.0-pro — Throughput: 205.3 tps, Latency: 75736.00s
  14. OpenAI gpt-5-nano — Throughput: 195.5 tps, Latency: 822.59s
  15. OpenAI o3 — Throughput: 184.6 tps, Latency: 0.00s
  16. Google Gemini 3.1 Pro Preview — Throughput: 179.1 tps, Latency: 22592.51s
  17. Google Gemini 3.5 Flash — Throughput: 168.2 tps, Latency: 18623.14s
  18. Google Gemini 3 Flash Preview — Throughput: 160.1 tps, Latency: 11859.89s
  19. xAI Grok 4.3 — Throughput: 158.4 tps, Latency: 12835.00s
  20. OpenAI GPT Image 1.5 — Throughput: 155.2 tps, Latency: 43711.00s
  21. Google gemini-2.5-flash — Throughput: 151.9 tps, Latency: 35168.00s
  22. OpenAI dall-e-3 — Throughput: 142.7 tps, Latency: 29253.50s
  23. xAI grok-3 — Throughput: 133.7 tps, Latency: 616.00s
  24. Anthropic Claude Claude Haiku 4.5 — Throughput: 130.6 tps, Latency: 4621.16s
  25. OpenAI gpt-4-gizmo-* — Throughput: 118.8 tps, Latency: 5586.06s
  26. Anthropic Claude Claude Sonnet 5 — Throughput: 113.9 tps, Latency: 7087.13s
  27. OpenAI GPT Image 2.5 Dev — Throughput: 109.0 tps, Latency: 44028.70s
  28. Z.ai GLM 5.2 — Throughput: 105.7 tps, Latency: 628.00s
  29. OpenAI gpt-5-mini — Throughput: 101.9 tps, Latency: 3849.00s
  30. Anthropic Claude Claude Opus 4.7 — Throughput: 98.4 tps, Latency: 3037.00s
  31. Z.ai GLM 5.1 — Throughput: 92.7 tps, Latency: 942.00s
  32. Google Gemini 3.1 Pro Preview Customtools — Throughput: 91.2 tps, Latency: 42673.33s
  33. Z.ai GLM 5.3 Flash — Throughput: 88.4 tps, Latency: 416.50s
  34. Google Gemini 3.7 Flash — Throughput: 88.0 tps, Latency: 11196.83s
  35. xAI Grok 4.20 — Throughput: 87.4 tps, Latency: 0.00s
  36. DeepSeek DeepSeek-V4-Flash — Throughput: 87.1 tps, Latency: 4568.90s
  37. Anthropic Claude Claude Opus 5 — Throughput: 86.8 tps, Latency: 3184.47s
  38. OpenAI GPT Image 2 Develop — Throughput: 81.7 tps, Latency: 25212.97s
  39. Google gemini-2.5-pro — Throughput: 81.2 tps, Latency: 50125.00s
  40. Alibaba qwen3-vl-plus — Throughput: 79.5 tps, Latency: 1297.00s
  41. xAI Grok 4.5 — Throughput: 78.7 tps, Latency: 11214.97s
  42. Google Gemini 3.8 Flash — Throughput: 78.4 tps, Latency: 12756.54s
  43. xAI grok-4-1-fast-reasoning — Throughput: 78.2 tps, Latency: 0.00s
  44. OpenAI gpt-image-1-mini — Throughput: 76.5 tps, Latency: 13478.00s
  45. DeepSeek DeepSeek-V4-Pro-0813 — Throughput: 76.0 tps, Latency: 1437.25s
  46. Z.ai GLM 5.3 — Throughput: 74.4 tps, Latency: 0.00s
  47. OpenAI Chat Latest (GPT-5.5 Instant) — Throughput: 72.8 tps, Latency: 0.00s
  48. Anthropic Claude Claude Opus 4.8 — Throughput: 70.1 tps, Latency: 3819.06s
  49. Thinking Machines Thinking Machines: Inkling — Throughput: 68.6 tps, Latency: 0.00s
  50. xAI Grok Build 0.1 — Throughput: 67.4 tps, Latency: 0.00s
  51. Anthropic Claude Claude Opus 4.6 — Throughput: 66.2 tps, Latency: 7395.29s
  52. DeepSeek DeepSeek-V4-Flash 0731 — Throughput: 65.2 tps, Latency: 1073.82s
  53. Google Gemini 3.1 Flash-Lite Preview — Throughput: 64.5 tps, Latency: 1027.67s
  54. Alibaba qwen3-235b-a22b-thinking-2507 — Throughput: 61.5 tps, Latency: 347.00s
  55. OpenAI GPT Image 2.5 Flare — Throughput: 60.9 tps, Latency: 23470.21s
  56. Moonshot Kimi K2.6 — Throughput: 60.2 tps, Latency: 964.00s
  57. Anthropic Claude Claude Sonnet 4.6 — Throughput: 59.2 tps, Latency: 5684.37s
  58. Z.ai GLM 5 — Throughput: 57.9 tps, Latency: 769.50s
  59. Moonshot kimi-k2-thinking — Throughput: 56.0 tps, Latency: 419.50s
  60. xAI Grok 4.6 — Throughput: 55.4 tps, Latency: 12215.67s
  61. MiniMax MiniMax-M2.7 Highspeed — Throughput: 54.8 tps, Latency: 0.00s
  62. OpenAI GPT-6 Astra — Throughput: 53.9 tps, Latency: 7807.80s
  63. DeepSeek DeepSeek-V3.2 — Throughput: 53.2 tps, Latency: 1238.86s
  64. Anthropic Claude Claude Fable 5 — Throughput: 53.0 tps, Latency: 14680.30s
  65. OpenAI gpt-3.5-turbo — Throughput: 52.8 tps, Latency: 0.00s
  66. Z.ai GLM 4.7 — Throughput: 50.3 tps, Latency: 1082.00s
  67. OpenAI GPT Image 2.5 Sunburst — Throughput: 47.7 tps, Latency: 31516.85s
  68. Anthropic Claude Claude Fable 5.1 — Throughput: 47.6 tps, Latency: 5663.91s
  69. OpenAI GPT-5.5 Pro — Throughput: 46.6 tps, Latency: 465.00s
  70. Anthropic Claude Claude Sonnet 4.5 — Throughput: 46.0 tps, Latency: 1833.95s
  71. MiniMax MiniMax M3 — Throughput: 46.0 tps, Latency: 1363.70s
  72. DeepSeek deepseek-v3.1 — Throughput: 45.2 tps, Latency: 2699.00s
  73. OpenAI gpt-4.1 — Throughput: 44.6 tps, Latency: 18569.09s
  74. MiniMax MiniMax-M2.1-lightning — Throughput: 44.4 tps, Latency: 0.00s
  75. OpenAI GPT-5.1 — Throughput: 43.6 tps, Latency: 14343.35s
  76. xAI grok-3-mini — Throughput: 43.1 tps, Latency: 0.00s
  77. xAI grok-4-fast-reasoning — Throughput: 42.5 tps, Latency: 0.00s
  78. Alibaba qwen-plus-latest — Throughput: 42.3 tps, Latency: 0.00s
  79. xAI grok-4-0709 — Throughput: 42.0 tps, Latency: 4728.44s
  80. OpenAI GPT-5.6 Luna — Throughput: 40.1 tps, Latency: 5637.80s
  81. Alibaba qvq-plus — Throughput: 40.0 tps, Latency: 487.00s
  82. xAI Grok 4.20 (Non-Reasoning) — Throughput: 38.6 tps, Latency: 1513.09s
  83. Google Nano Banana (Gemini 2.5 Flash Image 🍌) — Throughput: 38.2 tps, Latency: 6913.71s
  84. DeepSeek deepseek-v3 — Throughput: 37.7 tps, Latency: 2384.42s
  85. Google Gemini 3.5 Flash-Lite — Throughput: 37.5 tps, Latency: 4774.04s
  86. ByteDance Doubao-Seed-2.0-Code — Throughput: 36.2 tps, Latency: 0.00s
  87. OpenAI GPT Image 2 — Throughput: 35.4 tps, Latency: 34376.67s
  88. OpenAI GPT-5.6 Terra — Throughput: 34.5 tps, Latency: 58728.28s
  89. Google Gemini 3 Pro Preview — Throughput: 34.1 tps, Latency: 0.00s
  90. Moonshot Kimi K3 — Throughput: 33.8 tps, Latency: 4962.00s
  91. Google Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image 🍌) — Throughput: 33.4 tps, Latency: 44708.07s
  92. Xiaomi MiMo-V2.5 — Throughput: 31.6 tps, Latency: 17348.58s
  93. OpenAI GPT-5.6 Sol — Throughput: 31.5 tps, Latency: 4496.53s
  94. OpenAI GPT-5.2 — Throughput: 31.5 tps, Latency: 36212.29s
  95. OpenAI GPT-5.4 — Throughput: 30.3 tps, Latency: 12145.89s
  96. Alibaba qwen3-vl-flash — Throughput: 29.4 tps, Latency: 0.00s
  97. OpenAI GPT Image 1.0 — Throughput: 28.4 tps, Latency: 30244.50s
  98. OpenAI GPT-5.5 — Throughput: 28.4 tps, Latency: 1561.78s
  99. OpenAI gpt-4o — Throughput: 28.1 tps, Latency: 5627.43s
  100. ByteDance Doubao-Seed-2.0-mini — Throughput: 27.5 tps, Latency: 0.00s