Back to feed
Vercel
Vercel
7/21/2026
Service tiers now available on AI Gateway

Service tiers now available on AI Gateway

Short summary

Vercel's AI Gateway now supports service tiers (default, priority, flex) for OpenAI and Gemini models, letting developers optimize per-request for latency, throughput, or cost. Priority costs ~1.8-2x default for faster processing; flex costs ~0.5x for latency-tolerant jobs. Tiers work across all AI Gateway API formats and billing adjusts automatically based on the tier actually used.

  • Three tiers: default (baseline), priority (~1.8-2x cost, faster), flex (~0.5x cost, higher latency)
  • Works across OpenAI and Gemini models via providerOptions.gateway.serviceTier
  • Billing reflects actual tier used, with best-effort fallback to default if a tier can't be applied

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more