ToolsSeptember 18, 2026via AWS Machine Learning Blog
Introducing Amazon SageMaker HyperPod Inference Gateway
Why it matters
SageMaker HyperPod Inference Gateway is a production inference optimization tool that reduces latency through GPU-aware request routing on EKS — actionable for teams running inference at scale who want plug-and-play performance gains.
Key signals
- Kubernetes-native routing add-on for Amazon EKS
- GPU-aware request routing using real-time GPU signals
- Up to 82% reduction in first-token latency
- No changes required to model servers or client applications
- Part of Amazon SageMaker HyperPod family
- Amazon SageMaker HyperPod Inference Gateway launched
- Kubernetes-native, GPU-aware routing add-on for Amazon EKS
- Real-time GPU signal-based request routing
- Published September 18, 2026
The hook
82% faster first-token latency. Amazon just shipped a Kubernetes routing layer that makes inference servers smarter without touching your models.
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add-on for Amazon EKS. It uses real-time GPU signals to send each inference request to the best-suited pod, cutting first-token latency by up to 82% with no changes to your model servers or client applications.