ChipsThe story, in brief

Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Agentic AI is eating token meters whole. Enterprises running agents in production are hitting 10x consumption vs. chat — and per-token pricing can't scale it.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

Per-token pricing was designed for chat-style experimentation, but agentic workloads (multi-step reasoning, looping, tool calls) consume tokens at an order of magnitude higher rate. Enterprises need to model consumption economics and negotiate contract structures (fixed seats, compute units, time-based) before agent deployments go wide.

The key facts

5 to know
  1. Futurum report: 'The Off Ramp From Per-Token Pricing' sponsored by QumulusAI

  2. Key finding: per-token pricing optimal for enterprise experimentation but breaks under production agentic AI workloads

  3. Agentic systems consume dramatically higher token volumes than single-turn chat due to multi-step reasoning, tool loops, and internal reasoning chains

  4. Enterprises lack standardized consumption forecasting models for agent deployments

  5. Alternative pricing models emerging: fixed-seat licensing, compute-unit bundles, time-based consumption caps

Go to the source

SiliconAnglesiliconangle.com

Publisher excerpt: Per-token pricing was the best thing to happen to enterprises looking to experiment with artificial intelligence, but it may be the worst thing for AI in production. That’s the quandary at the center of a new Futurum report, “The Off Ramp From Per-Token Pricing,” sponsored by neocloud provider…
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips