ChipsThe story, in brief

DeepSeek overtakes Google on volume, cost per token falls 13.6%

DeepSeek just became the second-largest lab by token volume. Cost per token fell 13.6% in July alone. Here's what actually happened in enterprise AI.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Market-shifting data: Open-weight models (DeepSeek, Moonshot, Z.ai) are capturing production volume at commodity pricing, forcing frontier labs into a two-tier market. Anthropic holds 65% of spend on 30% of volume by pricing power alone; cheaper alternatives exist but require task-switching. This is the first month where open-weight models meaningfully competed on capability-per-dollar, not just cost.

The key facts

11 to know
  1. DeepSeek captured 25% of gateway token volume in July, surpassing Google's 11% (down from Google's 40% in April)

  2. DeepSeek V4 Flash ran 19% of all tokens on the gateway, 70% more than the next model

  3. Open-weight models' share of gateway spend more than doubled to 8.6% in July, driven by Kimi K3 and GLM 5.2

  4. Average price per token fell 13.6% in July despite 37% spend growth, driven by model mix shift not price cuts

  5. Anthropic collected 65.1% of gateway spend on 30% of token volume at 4.4x average price per token

  6. Kimi K3 (released July 16) scaled to 8th place by volume within two weeks, tripled daily volume by month end

  7. 81% of July tokens ran on models not present on gateway six months prior

  8. Three in four teams with >10M monthly tokens changed their model mix by at least 10%; half changed by 25%+

  9. Anthropic's Haiku 4.5 (cheapest offering) costs 68% of gateway average; DeepSeek V4 Flash costs 6% of that average

  10. In coding agents (largest workload): DeepSeek ran 32% of volume, Anthropic collected 80%+ of spend

  11. Moonshot's share of total gateway spend quadrupled to 2.3% in July

Go to the source

Vercel Blogvercel.com

Publisher excerpt: AI Gateway Production Index — August 2026 Every month, routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Production Index reports from , ,…
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips