Live model performance metrics accessible via AI Gateway
Vercel's AI Gateway now surfaces live performance metrics across 100+ models—helping devs pick winners based on real throughput and latency data, not benchmarks.

Why it matters
Developers now have transparency into actual model performance (latency, throughput) across providers in production. This shifts model selection from benchmark claims to observable, real-world metrics—lowering friction in the model-picking workflow.
The key facts
12 to knowVercel AI Gateway now displays P50 and P95 latency (TTFT) and throughput metrics
Metrics updated hourly based on live customer requests
Sortable by latency (time-to-first-token) and throughput (tokens/second)
Metrics available on model list pages, individual model detail pages, and via REST API
Provider-level performance breakdowns enable cross-provider comparison for same model
Data includes both P50 (best case) and P95 (tail latency) percentiles
Metrics updated hourly across hundreds of models
P50 latency and throughput visible per model and per provider
REST API access to live performance data (P50/P95)
Sortable columns for latency (time-to-first-token) and throughput (tokens/sec)
Data pulled from live AI Gateway customer requests
Metrics available via three interfaces: model list, detail pages, and REST endpoints
Go to the source
Vercel Blogvercel.com
Publisher excerpt: AI Gateway now displays throughput and latency metrics across hundreds of models, helping you choose the right model based on live performance data. Metrics appear in three places and are updated every hour: The AI Gateway now includes sortable columns for latency and throughput. Each row displays…