Codeseys/composer-replication-framework
0
1---2title: Amazon SageMaker HyperPod - Amazon SageMaker AI3id: amazon-sagemaker-hyperpod-amazon-sagemaker-ai4tags:5- socratic-mcts-swe-worldmodel-8f6dea6created: '2026-06-09T04:24:49.182604Z'7source: https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod.html8source_domain: docs.aws.amazon.com9fetched_at: '2026-06-09T04:24:46.875226Z'10fetch_provider: builtin11status: draft12type: note13deprecated: false14summary: Amazon SageMaker HyperPod - Amazon SageMaker AI15---16 17Amazon SageMaker HyperPod - Amazon SageMaker AI18View a markdown version of this page19Amazon SageMaker HyperPod - Amazon SageMaker AI20Documentation21Amazon SageMaker22Developer Guide23AWS Regions supported by SageMaker HyperPod24Amazon SageMaker HyperPod25SageMaker HyperPod helps you provision resilient clusters for running machine learning (ML)26 workloads and developing state-of-the-art models such as large language models (LLMs),27 diffusion models, and foundation models (FMs). It accelerates development of FMs by removing28 undifferentiated heavy-lifting involved in building and maintaining large-scale compute29 clusters powered by thousands of accelerators such as AWS Trainium and NVIDIA A100 and30 H100 Graphical Processing Units (GPUs). When accelerators fail, the resiliency features of31 SageMaker HyperPod monitor the cluster instances automatically detect and replace the faulty32 hardware on the fly so that you can focus on running ML workloads.33To get started, check34Prerequisites for using SageMaker HyperPod35, set up36AWS Identity and Access Management for SageMaker HyperPod37, and choose one of the following38 orchestrator options supported by SageMaker HyperPod.39Slurm support in SageMaker HyperPod40SageMaker HyperPod provides support for running machine learning workloads on resilient clusters41 by integrating with Slurm, an open-source workload manager. Slurm support in SageMaker HyperPod42 enables seamless cluster orchestration through Slurm cluster configuration, allowing you to43 set up head, login, and worker nodes on the SageMaker HyperPod clusters This integration also44 facilitates Slurm-based job scheduling for running ML workloads on the cluster, as well as45 direct access to cluster nodes for job scheduling. With HyperPod's lifecycle46 configuration support, you can customize the computing environment of the clusters to meet47 your specific requirements. Additionally, by leveraging the Amazon SageMaker AI distributed training48 libraries, you can optimize the clusters' performance on AWS computing and network49 resources. To learn more, see50Orchestrating SageMaker HyperPod clusters with Slurm51.52Amazon EKS support in SageMaker HyperPod53SageMaker HyperPod also integrates with Amazon EKS to enable large-scale training of foundation54 models on long-running and resilient compute clusters. This allows cluster admin users to55 provision HyperPod clusters and attach them to an EKS control plane, enabling56 dynamic capacity management, direct access to cluster instances, and resiliency57 capabilities. For data scientists, Amazon EKS support in HyperPod allows running58 containerized workloads for training foundation models, inference on the EKS cluster, and59 leveraging the job auto-resume capability for Kubeflow PyTorch training. The architecture60 involves a 1-to-1 mapping between an EKS cluster (control plane) and a HyperPod61 cluster (worker nodes) within a VPC, providing a tightly integrated solution for running62 large-scale ML workloads. To learn more, see63Orchestrating SageMaker HyperPod clusters with Amazon EKS64.65UltraServers with HyperPod66HyperPod with UltraServers delivers AI computing power by integrating67 NVIDIA superchips into a cohesive, high-performance infrastructure. Each NVL72 UltraServer 68 combines 18 instances with 72 NVIDIA Blackwell GPUs interconnected via NVLink, enabling69 faster inference and faster training performance compared to previous generation instances. This 70 architecture is particularly valuable for organizations working with trillion-parameter foundation 71 models, as the unified GPU memory allows entire models to remain within a single 72 NVLink domain, eliminating cross-node networking bottlenecks. HyperPod 73 enhances this hardware advantage74 with intelligent topology-aware scheduling that optimizes workload placement, automatic instance 75 replacement to minimize disruptions, and flexible deployment options that support both dedicated and 76 shared resource configurations. For teams pushing the boundaries of model size and performance, this77 integration provides the computational foundation needed to train and deploy the most advanced AI 78 models with unprecedented efficiency.79SageMaker HyperPod automatically optimizes instance placement across your UltraServers. 80 By default, HyperPod prioritizes all instances in one UltraServer before using a different one. 81 For example, if you want 14 instances and have 2 UltraServers in your plan, SageMaker AI uses all of the 82 instances in the first UltraServer. If you want 20 instances, SageMaker AI uses all 18 instances in the first83 UltraServer and then uses 2 more from the second.84AWS Regions supported by SageMaker HyperPod85SageMaker HyperPod is available in the following AWS Regions.86us-east-187us-east-288us-west-189us-west-290eu-central-191eu-north-192eu-west-193eu-west-294eu-south-295ap-south-196ap-southeast-197ap-southeast-298ap-southeast-399ap-southeast-4100ap-northeast-1101ap-northeast-2102sa-east-1103Topics104Amazon SageMaker HyperPod quickstart105Prerequisites for using SageMaker HyperPod106AWS Identity and Access Management for SageMaker HyperPod107Customer managed AWS KMS key encryption for SageMaker HyperPod108SageMaker HyperPod recipes109Orchestrating SageMaker HyperPod clusters with Slurm110Orchestrating SageMaker HyperPod clusters with Amazon EKS111Using topology-aware scheduling in Amazon SageMaker HyperPod112Deploying models on Amazon SageMaker HyperPod113HyperPod in Studio114SageMaker HyperPod references115Amazon SageMaker HyperPod release notes116Amazon SageMaker HyperPod AMI117Javascript is disabled or is unavailable in your browser.118To use the Amazon Web Services Documentation, Javascript must be enabled. Please refer to your browser's Help pages for instructions.119Document Conventions120Custom images121Quickstart122Did this page help you? - Yes123Thanks for letting us know we're doing a good job!124If you've got a moment, please tell us what we did right so we can do more of it.125Did this page help you? - No126Thanks for letting us know this page needs work. We're sorry we let you down.127If you've got a moment, please tell us how we can make the documentation better.