Accelerate Generative AI Inference on Amazon SageMaker AI with G7e Instances
NVIDIA RTX PRO 6000 Blackwell just landed on SageMaker. Single-node inference for 120B models just got cheaper.

Why it matters
AWS is lowering the barrier to deploy large foundation models by offering cost-effective GPU infrastructure via SageMaker G7e instances. This matters because it directly reduces capex friction for enterprises running inference at scale.
The key facts
5 to knowG7e instances powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs now available on Amazon SageMaker AI
Each GPU provides 96 GB of GDDR7 memory
Configurable node counts: 1, 2, 4, and 8 GPU instances
G7e.2xlarge (single-node) supports models like GPT-OSS-120B, Nemotron-3-Super-120B, Qwen3.5-35B
Positions as cost-effective alternative for foundation model inference
Go to the source
AWS Machine Learning Blogaws.amazon.com
Publisher excerpt: Today, we are thrilled to announce the availability of G7e instances powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs on Amazon SageMaker AI. You can provision nodes with 1, 2, 4, and 8 RTX PRO 6000 GPU instances, with each GPU providing 96 GB of GDDR7 memory. This launch provides the…