GLM 5.2 Fast via Wafer now available on AI Gateway
2x throughput. Vercel's AI Gateway now serves GLM-5.2 Fast via Wafer with 170+ tok/s on small context—fastest serverless option for sustained generation.

Why it matters
Vercel is expanding its AI Gateway infrastructure to include optimized routing for Zhipu's GLM-5.2 model, offering developers faster inference at no platform markup. This is a competitive move in the app-layer tooling space where latency and cost directly impact product UX.
The key facts
13 to knowGLM-5.2 Fast available on Vercel AI Gateway via Wafer
2x higher throughput vs other serverless providers
Small context: 170+ tok/s
Large context: 200+ tok/s
No platform fee on inference, no data retention, BYOK support
Unified API for usage tracking, retries, failover, performance optimization
GLM-5.2 Fast via Wafer now available on Vercel AI Gateway
2x higher throughput vs. other serverless providers for GLM-5.2
Small context: 170+ tokens/second
Large context: 200+ tokens/second
Zero platform fee on inference, including BYOK requests
Provider pricing reflected with no markup
Built-in custom reporting, Zero Data Retention support, budget controls
Go to the source
Vercel Blogvercel.com
Publisher excerpt: GLM 5.2 Fast via Wafer is now available on .AI Gateway Based on our own benchmarking across small-context, large-context, and tool-call scenarios, Wafer delivers a 2x higher throughput than other providers serving GLM-5.2 on serverless, leading on decode and end-to-end speed for sustained…