On August 24, 2026, Amazon SageMaker HyperPod introduced managed Ray support on Amazon Elastic Kubernetes Service (EKS), enabling users to create and monitor Ray clusters directly from the platform. Announced on the AWS Machine Learning Blog, this new capability integrates Ray, an open-source framework for distributed Python workloads, with HyperPod's purpose-built infrastructure for large-scale machine learning.
Simplified Cluster Management
Until now, running Ray on Kubernetes required data scientists to manage several complex tasks, including writing YAML manifests, handling Docker image rebuilds, and configuring observability tools. With the new managed Ray support, these tasks are automated, allowing users to focus on their machine learning workloads.
Enhanced Observability and Integration
Users can now connect JupyterLab and Code Editor notebooks to live Ray clusters, benefiting from out-of-the-box observability. This integration simplifies the process of running resilient distributed training and accelerated inference directly from SageMaker Studio.
Built on Open-Source and Standard APIs
The solution leverages the open-source KubeRay project and standard Ray APIs, ensuring compatibility and ease of use. This approach allows data scientists to take full advantage of Ray's capabilities without the overhead of manual configuration.
Key Benefits
Simplified Cluster Management: Automated management of Ray clusters on Amazon EKS.
Seamless Integration: Direct connection of JupyterLab and Code Editor notebooks to Ray clusters.
Out-of-the-Box Observability: Built-in monitoring and observability tools.
Accelerated Machine Learning Workflows: Enhanced capabilities for distributed training and inference.
This launch marks a significant step forward in simplifying the deployment and management of Ray clusters on Amazon EKS, empowering data scientists to focus more on their machine learning tasks.