ToolsThe story, in brief

AI Gateway is now generally available

Not a pilot. Vercel's AI Gateway now routes to 100+ models with sub-20ms latency and zero token markup.

Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
New tools for building and creating with AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Vercel democratizes multi-model inference by removing pricing friction and operational complexity. For founders building AI products, this eliminates vendor lock-in while simplifying cost tracking and failover—critical for production deployments.

The key facts

7 to know
  1. AI Gateway now generally available

  2. Sub-20ms latency routing across multiple inference providers

  3. Transparent pricing with no markup on tokens

  4. Supports Bring Your Own Keys (BYOK)

  5. Automatic failover for higher availability

  6. Works with Vercel AI SDK and OpenAI-compatible endpoints

  7. Single model string switch for implementation

Go to the source

Vercel Blogvercel.com

Publisher excerpt: is now generally available, providing a single unified API to access hundreds of AI models with transparent pricing and built-in observability.AI Gateway With sub-20ms latency routing across multiple inference providers, AI Gateway delivers: You can use AI Gateway with the or through the…
Read original report
Back to today's editionMore tools news

The wider picture

View all
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Tools01

Vast Data launches confidential computing service for sensitive workloads

A new product feature addresses a real blocker for enterprise AI adoption — the need to run models on confidential data while protecting both the data and model IP. DataEnclave using Nvidia's confidential computing tech lowers friction for regulated industries.

SiliconAngle
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools02

Transformers now runs llama.cpp quants

Hugging Face's mainstream library cutting friction between open-weight models and llama.cpp's efficient inference stack. Practitioners can now load and run quants in production workflows without format bridging or custom code.

Hugging Face Blog
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools03

Meta’s Muse is outpacing ChatGPT’s early mobile launch

Meta's new AI agent is proving consumer adoption curves have compressed; mobile AI products are now table stakes, and Muse's early traction signals that distribution (not capability) drives initial user growth in a crowded market.

TechCrunch AI