r/MachineLearning
7/23/2026
![Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]](https://preview.redd.it/7ejtvp684xeh1.png?width=140&height=65&auto=webp&s=10790ba444afd733ece8c54a8b9da99969a86066)
Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]
Short summary
A benchmark of 10 realistic product tasks across GPT, Claude, Gemini, and Kimi APIs reveals a 10.6x total cost spread despite published rates differing by only 2x. The gap is driven by invisible reasoning/thinking tokens billed at output rates but never shown in responses. Findings align with CostBench (ACL 2026) research showing models routinely fail to choose cost-optimal plans.
- •10.6x cost spread across 4 LLM providers on identical tasks despite only 2x published price difference
- •Invisible reasoning tokens billed at output rate are the primary cost driver
- •Full methodology and prompts are open-source on GitHub
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



