ChipsThe story, in brief

Cloudflare Builds High-Performance Infrastructure for Running LLMs

Cloudflare just separated LLM inference into two optimized pipelines. Here's why that matters for your compute costs.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Cloudflare is building specialized infrastructure to reduce the hardware and operational cost of running LLMs at scale. By decoupling input processing from output generation, they're addressing a real constraint for companies deploying models globally—and positioning themselves as an alternative to traditional cloud providers for AI workloads.

The key facts

8 to know
  1. Cloudflare announces LLM-optimized infrastructure

  2. Architecture separates model input processing and output generation into distinct systems

  3. Infrastructure deployed across Cloudflare's global network

  4. Focus on reducing hardware costs and handling high text throughput

  5. Cloudflare announced new infrastructure for running LLMs globally

  6. Architecture separates model input processing and output generation onto different optimized systems

  7. Design addresses hardware costs and high-volume text I/O handling

  8. Infrastructure spans Cloudflare's global network

Go to the source

InfoQ AI/MLinfoq.com

Publisher excerpt: Cloudflare has recently announced new infrastructure designed to run large AI language models across its global network. As these models rely on costly hardware and must handle large volumes of incoming and outgoing text, Cloudflare separated the model's input processing and output generation onto…
Read original report
Back to today's editionMore chips news

The wider picture

View all
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips01

Google Adds Cycle-Level Kernel Profiling to XProf

A developer-facing tooling improvement that directly enables better TPU utilization and kernel optimization. Practitioners building custom Pallas kernels can now see exactly where cycles are spent, shifting from guesswork to data-driven tuning.

InfoQ AI/ML
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips02

Civo unveils first of 40 planned edge data center sites across UK

Edge compute infrastructure is becoming critical for low-latency AI inference and agentic workloads. Civo's distributed network strategy reflects growing demand for regional AI compute capacity outside centralized cloud zones — a structural shift in how AI workloads are deployed.

ITPro
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips03

China reviews dependence on Broadcom switches in data centres

China is auditing its reliance on foreign networking hardware for AI data centers as part of a broader push to build domestic alternatives. This reshapes global compute buildout economics and chip supply chains at a moment when AI capacity is the competitive moat.

Financial Times Technology