ToolsThe story, in brief

GLM 5.2 Fast via Wafer now available on AI Gateway

2x throughput. Vercel's AI Gateway now serves GLM-5.2 Fast via Wafer with 170+ tok/s on small context—fastest serverless option for sustained generation.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Vercel is expanding its AI Gateway infrastructure to include optimized routing for Zhipu's GLM-5.2 model, offering developers faster inference at no platform markup. This is a competitive move in the app-layer tooling space where latency and cost directly impact product UX.

The key facts

13 to know
  1. GLM-5.2 Fast available on Vercel AI Gateway via Wafer

  2. 2x higher throughput vs other serverless providers

  3. Small context: 170+ tok/s

  4. Large context: 200+ tok/s

  5. No platform fee on inference, no data retention, BYOK support

  6. Unified API for usage tracking, retries, failover, performance optimization

  7. GLM-5.2 Fast via Wafer now available on Vercel AI Gateway

  8. 2x higher throughput vs. other serverless providers for GLM-5.2

  9. Small context: 170+ tokens/second

  10. Large context: 200+ tokens/second

  11. Zero platform fee on inference, including BYOK requests

  12. Provider pricing reflected with no markup

  13. Built-in custom reporting, Zero Data Retention support, budget controls

Go to the source

Vercel Blogvercel.com

Publisher excerpt: GLM 5.2 Fast via Wafer is now available on .AI Gateway Based on our own benchmarking across small-context, large-context, and tool-call scenarios, Wafer delivers a 2x higher throughput than other providers serving GLM-5.2 on serverless, leading on decode and end-to-end speed for sustained…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools