Vercel
7/21/2026

Service tiers now available on AI Gateway
Short summary
Vercel's AI Gateway now supports service tiers (default, priority, flex) for OpenAI and Gemini models, letting developers optimize per-request for latency, throughput, or cost. Priority costs ~1.8-2x default for faster processing; flex costs ~0.5x for latency-tolerant jobs. Tiers work across all AI Gateway API formats and billing adjusts automatically based on the tier actually used.
- •Three tiers: default (baseline), priority (~1.8-2x cost, faster), flex (~0.5x cost, higher latency)
- •Works across OpenAI and Gemini models via providerOptions.gateway.serviceTier
- •Billing reflects actual tier used, with best-effort fallback to default if a tier can't be applied
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



