ChipsThe story, in brief

NVIDIA and Google infrastructure cuts AI inference costs

10x cheaper. Google and NVIDIA just solved AI inference's biggest cost problem with bare-metal A5X instances.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

Google and NVIDIA are directly addressing inference economics—the bottleneck blocking AI deployment at scale. Hardware-software codesign targeting 10x cost reduction changes the unit economics of production AI workloads.

The key facts

5 to know
  1. A5X bare-metal instances announced

  2. Built on NVIDIA Vera Rubin NVL72 rack-scale systems

  3. Claims up to 10x lower inference costs

  4. Hardware-software codesign approach

  5. Announced at Google Cloud Next 2026

Go to the source

AI Newsartificialintelligence-news.com

Publisher excerpt: At the Google Cloud Next conference, Google and NVIDIA outlined their hardware roadmap designed to address the cost of AI inference at scale. The companies detailed the new A5X bare-metal instances, which run on NVIDIA Vera Rubin NVL72 rack-scale systems. Through hardware and software codesign,…
Read original report
Back to today's editionMore chips news

Keep reading

Related stories

More from Chips