Back to feed
r/MachineLearning
r/MachineLearning
6/24/2026
I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]

I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]

Short summary

Author compiled LLM inference pricing from 7+ providers and discovered caching policies vary dramatically—sometimes 10x+ cheaper on cache hits than misses. For agents, RAG pipelines, and multi-turn conversations, caching strategy matters far more than headline token prices. Includes spreadsheet comparing token costs, context windows, and provider availability.

  • Caching costs vary 10x+ across providers for same model (DeepSeek V4 Pro example)
  • Caching matters more than token price for agents and RAG pipelines
  • Spreadsheet compares 7+ providers (OpenRouter, DeepSeek, Together, Fireworks, Groq) on pricing, context windows, and availability

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more