Dev.to
8/3/2026

Why AI price cuts of 80% still leave your cloud bill rising: the infrastructure cost shift
Original: GPT-5.6 Luna Just Cut Prices 80% Your AI Bill Is Still Going Up, and Here’s the Math
Short summary
AI model token prices are dropping fast, but total AI infrastructure bills keep rising due to the Jevons paradox: cheaper tokens unlock 10× more usage. The real cost shift is from ~80% tokens in 2024 toward ~70% runtime, memory, observability, and GPU costs — infrastructure that follows traditional cloud pricing curves, not model price wars. Teams need standard FinOps discipline: schedule idle agent runtimes, tune observability retention, route requests to cheaper model tiers, and add anomaly detection for AI resources.
- •Token price cuts drive usage up 10×, netting higher bills (Jevons paradox)
- •AI spend shifting from ~80% tokens to ~70% runtime/memory/observability/GPU
- •Standard FinOps levers (scheduling, rightsizing, retention tuning) now apply to AI infrastructure
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



