ToolsThe story, in brief

Introducing Active CPU pricing for Fluid compute

Up to 85% cost cuts. Vercel's new Active CPU pricing charges only when your inference code actually runs—not idle time.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

As AI inference workloads scale, serverless pricing models that charge for idle time become prohibitively expensive. Vercel's Active CPU pricing directly addresses the economics of long-running AI agents and inference backends, making real-time AI deployments more cost-competitive.

The key facts

11 to know
  1. Vercel launches Active CPU pricing for Fluid compute

  2. Up to 85% cost savings claimed through optimizations like in-function concurrency

  3. Targets I/O-bound workloads: AI inference, agents, MCP servers

  4. Pay-per-active-CPU model (not idle-time billing)

  5. Designed for long-running, unpredictable workloads

  6. Published June 25, 2025

  7. Up to 85% cost savings through optimizations like in-function concurrency

  8. New pricing model: pay CPU rates only when code actively uses CPU

  9. Targets I/O bound backends: AI inference, agents, MCP servers

  10. Addresses long-running, unpredictable workloads with frequent idle periods

  11. Fluid Compute became default model on Vercel

Go to the source

Vercel Blogvercel.com

Publisher excerpt: exists for a new class of workloads. I/O bound backends like AI inference, agents, MCP servers, and anything that needs to scale instantly, but often remains idle between operations. These workloads do not follow traditional, quick request-response patterns. They’re long-running, unpredictable, and…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools