ToolsThe story, in brief

Moonshot AI's Kimi K2 model is now supported in Vercel AI Gateway

Vercel just added Moonshot's Kimi K2 to AI Gateway. One API call. Five provider backends. No vendor lock-in.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Vercel AI Gateway expands model access to include Moonshot's mixture-of-experts model, enabling developers to route inference across multiple providers (Groq, DeepInfra, Fireworks, Parasail) with unified observability and failover—lowering switching costs and latency for production AI apps.

The key facts

11 to know
  1. Moonshot AI's Kimi K2 (MoE model) now integrated into Vercel AI Gateway

  2. Unified API support via AI SDK v5 with single model string update (moonshotai/kimi-k2)

  3. Multi-provider routing: Direct Moonshot AI, Groq, DeepInfra, Fireworks AI, Parasail

  4. Built-in observability, BYOK support, intelligent provider routing with auto-retries

  5. No additional provider accounts required for Gateway users

  6. Published July 15, 2025

  7. Kimi K2 (mixture-of-experts model) now available via Vercel AI Gateway

  8. Unified API across multiple providers: Moonshot AI, Groq, DeepInfra, Fireworks AI, Parasail

  9. Single model string update required ('moonshotai/kimi-k2')

  10. AI SDK v5 integration

  11. Eliminates need for separate provider accounts

Go to the source

Vercel Blogvercel.com

Publisher excerpt: You can now access , a new mixture-of-experts (MoE) language model from , using Vercel's with no other provider accounts required.Kimi K2Moonshot AIAI Gateway AI Gateway lets you call the model with a consistent unified API and just a single string update, track usage and cost, and configure…
Read original report
Back to today's editionMore tools news

The wider picture

View all
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Tools01

Vast Data launches confidential computing service for sensitive workloads

A new product feature addresses a real blocker for enterprise AI adoption — the need to run models on confidential data while protecting both the data and model IP. DataEnclave using Nvidia's confidential computing tech lowers friction for regulated industries.

SiliconAngle
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools02

Transformers now runs llama.cpp quants

Hugging Face's mainstream library cutting friction between open-weight models and llama.cpp's efficient inference stack. Practitioners can now load and run quants in production workflows without format bridging or custom code.

Hugging Face Blog
Paper-cut illustration of a coral software window opening into a three-dimensional drafting space.
AI illustration by KeyNews
Tools03

Meta’s Muse is outpacing ChatGPT’s early mobile launch

Meta's new AI agent is proving consumer adoption curves have compressed; mobile AI products are now table stakes, and Muse's early traction signals that distribution (not capability) drives initial user growth in a crowded market.

TechCrunch AI