ToolsThe story, in brief

Lower pricing with Active CPU pricing for Fluid compute

Not a pilot. Vercel just cut LLM inference costs in half with Active CPU pricing.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Vercel's new Active CPU pricing model for Fluid Compute directly reduces operational spend for AI workloads by eliminating idle-time billing—critical economics for cost-conscious teams running LLM inference and long-running agents at scale.

The key facts

10 to know
  1. Active CPU pricing reduces example workload cost from $0.31842/hour to $0.149/hour (~53% reduction)

  2. Pricing structure: $0.128/hour Active CPU + $0.0106/GB-hour Provisioned Memory + per-invocation charges

  3. Enabled by default for Hobby, Pro, and new Enterprise teams; existing Enterprise availability varies

  4. Targets LLM inference, AI agents, and other workloads with significant idle time

  5. Changes take effect after redeploy

  6. Active CPU pricing reduces costs from $0.31842/hour to $0.149/hour for Standard machine at 100% utilization (53% savings)

  7. Charges only for active CPU time ($0.128/hour), not idle periods

  8. Provisioned memory billed separately at $0.0106 per GB-hour

  9. Enabled by default for Hobby, Pro, and new Enterprise teams

  10. Use case: LLM inference, long-running AI agents, workloads with significant idle time

Go to the source

Vercel Blogvercel.com

Publisher excerpt: Vercel Functions on Fluid Compute now use Active CPU pricing, which charges for CPU only while it is actively doing work. This eliminates costs during idle time and reduces spend for workloads like LLM inference, long-running AI agents, or any task with idle time. Active CPU pricing is built on…
Read original report
Back to today's editionMore tools news

Keep reading

Related stories

More from Tools