Back to feed
Dev.to
Dev.to
6/16/2026
I measure how fast 42 LLMs actually answer. Here's the honest method.

I measure how fast 42 LLMs actually answer. Here's the honest method.

Short summary

Anton Gulin benchmarks 42 Ollama Cloud models on two key metrics: time to first token (response latency) and tokens per second (generation speed). His methodology uses fixed prompts, continuous retesting every 10 minutes, and reveals an 80x difference in response latency, with a smaller 30B model outperforming larger alternatives. The independent tracker at ollamatps.com is reproducible and updates live.

  • Independent tracker benchmarks 42 LLMs on response latency (TTFT) and generation speed (TPS)
  • Results reveal 80x variation in response time and counter-intuitive performance (smallest model is fastest)
  • Methodology is reproducible with fixed prompts, continuous retesting, and live updates at ollamatps.com

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more