Realtime voice, speech, and transcription now supported on AI Gateway
Realtime voice agents just went live on Vercel's AI Gateway. No markup. No platform fees. Same controls as text.

Why it matters
Vercel democratizes voice AI infrastructure by adding realtime voice, speech, and transcription to AI Gateway with cost parity to text models—lowering the barrier for developers to build conversational agents without vendor lock-in or hidden fees.
The key facts
9 to knowRealtime voice agents now supported on Vercel AI Gateway
Speech-to-text and text-to-speech capabilities in beta
Same observability and spend controls as text/image/video models
No markup or platform fees for voice models
Bring-your-own-key support included
Near real-time audio I/O for conversational agents
Tool-calling capability mid-conversation for voice agents
Browser-based playground for no-code testing
useRealtime hook for WebSocket management and audio capture
Go to the source
Vercel Blogvercel.com
Publisher excerpt: now supports voice and audio models. You can build realtime voice agents, generate speech from text, and transcribe audio to text. This provides the same observability, spend controls, and bring-your-own-key support as text, image, and video models in AI Gateway, with no markup or platform fees.…