Back to feed
r/MachineLearning
r/MachineLearning
7/23/2026
Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]

Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]

Short summary

A benchmark of 10 realistic product tasks across GPT, Claude, Gemini, and Kimi APIs reveals a 10.6x total cost spread despite published rates differing by only 2x. The gap is driven by invisible reasoning/thinking tokens billed at output rates but never shown in responses. Findings align with CostBench (ACL 2026) research showing models routinely fail to choose cost-optimal plans.

  • 10.6x cost spread across 4 LLM providers on identical tasks despite only 2x published price difference
  • Invisible reasoning tokens billed at output rate are the primary cost driver
  • Full methodology and prompts are open-source on GitHub

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more