r/MachineLearning
6/24/2026
![I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]](https://preview.redd.it/4vj50mvhu79h1.png?width=140&height=63&auto=webp&s=f53b566e7aa9a25215aa77fcf3ed0b16e426e2a1)
I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]
Short summary
Author compiled LLM inference pricing from 7+ providers and discovered caching policies vary dramatically—sometimes 10x+ cheaper on cache hits than misses. For agents, RAG pipelines, and multi-turn conversations, caching strategy matters far more than headline token prices. Includes spreadsheet comparing token costs, context windows, and provider availability.
- •Caching costs vary 10x+ across providers for same model (DeepSeek V4 Pro example)
- •Caching matters more than token price for agents and RAG pipelines
- •Spreadsheet compares 7+ providers (OpenRouter, DeepSeek, Together, Fireworks, Groq) on pricing, context windows, and availability
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



