ChipsThe story, in brief

Accelerate Generative AI Inference on Amazon SageMaker AI with G7e Instances

NVIDIA RTX PRO 6000 Blackwell just landed on SageMaker. Single-node inference for 120B models just got cheaper.

Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
The infrastructure powering AI.AI illustration by KeyNews
The KeyNews take

Why it matters

AWS is lowering the barrier to deploy large foundation models by offering cost-effective GPU infrastructure via SageMaker G7e instances. This matters because it directly reduces capex friction for enterprises running inference at scale.

The key facts

5 to know
  1. G7e instances powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs now available on Amazon SageMaker AI

  2. Each GPU provides 96 GB of GDDR7 memory

  3. Configurable node counts: 1, 2, 4, and 8 GPU instances

  4. G7e.2xlarge (single-node) supports models like GPT-OSS-120B, Nemotron-3-Super-120B, Qwen3.5-35B

  5. Positions as cost-effective alternative for foundation model inference

Go to the source

AWS Machine Learning Blogaws.amazon.com

Publisher excerpt: Today, we are thrilled to announce the availability of G7e instances powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs on Amazon SageMaker AI. You can provision nodes with 1, 2, 4, and 8 RTX PRO 6000 GPU instances, with each GPU providing 96 GB of GDDR7 memory. This launch provides the…
Read original report
Back to today's editionMore chips news

The wider picture

View all
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips01

Google Adds Cycle-Level Kernel Profiling to XProf

A developer-facing tooling improvement that directly enables better TPU utilization and kernel optimization. Practitioners building custom Pallas kernels can now see exactly where cycles are spent, shifting from guesswork to data-driven tuning.

InfoQ AI/ML
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips02

Civo unveils first of 40 planned edge data center sites across UK

Edge compute infrastructure is becoming critical for low-latency AI inference and agentic workloads. Civo's distributed network strategy reflects growing demand for regional AI compute capacity outside centralized cloud zones — a structural shift in how AI workloads are deployed.

ITPro
Paper-cut illustration of an amber microchip with circuit paths extending into a row of data-center cabinets.
AI illustration by KeyNews
Chips03

China reviews dependence on Broadcom switches in data centres

China is auditing its reliance on foreign networking hardware for AI data centers as part of a broader push to build domestic alternatives. This reshapes global compute buildout economics and chip supply chains at a moment when AI capacity is the competitive moat.

Financial Times Technology