Team Ai
Modelpublic

Codeseys/composer-replication-framework

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes
amazon-sagemaker-hyperpod-amazon-sagemaker-ai.md127 linesDownload Raw Back to notes
1---2title: Amazon SageMaker HyperPod - Amazon SageMaker AI3id: amazon-sagemaker-hyperpod-amazon-sagemaker-ai4tags:5- socratic-mcts-swe-worldmodel-8f6dea6created: '2026-06-09T04:24:49.182604Z'7source: https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod.html8source_domain: docs.aws.amazon.com9fetched_at: '2026-06-09T04:24:46.875226Z'10fetch_provider: builtin11status: draft12type: note13deprecated: false14summary: Amazon SageMaker HyperPod - Amazon SageMaker AI15---16 17Amazon SageMaker HyperPod - Amazon SageMaker AI18View a markdown version of this page19Amazon SageMaker HyperPod - Amazon SageMaker AI20Documentation21Amazon SageMaker22Developer Guide23AWS Regions supported by SageMaker HyperPod24Amazon SageMaker HyperPod25SageMaker HyperPod helps you provision resilient clusters for running machine learning (ML)26        workloads and developing state-of-the-art models such as large language models (LLMs),27        diffusion models, and foundation models (FMs). It accelerates development of FMs by removing28        undifferentiated heavy-lifting involved in building and maintaining large-scale compute29        clusters powered by thousands of accelerators such as AWS Trainium and NVIDIA A100 and30        H100 Graphical Processing Units (GPUs). When accelerators fail, the resiliency features of31        SageMaker HyperPod monitor the cluster instances automatically detect and replace the faulty32        hardware on the fly so that you can focus on running ML workloads.33To get started, check34Prerequisites for using SageMaker HyperPod35, set up36AWS Identity and Access Management for SageMaker HyperPod37, and choose one of the following38        orchestrator options supported by SageMaker HyperPod.39Slurm support in SageMaker HyperPod40SageMaker HyperPod provides support for running machine learning workloads on resilient clusters41        by integrating with Slurm, an open-source workload manager. Slurm support in SageMaker HyperPod42        enables seamless cluster orchestration through Slurm cluster configuration, allowing you to43        set up head, login, and worker nodes on the SageMaker HyperPod clusters This integration also44        facilitates Slurm-based job scheduling for running ML workloads on the cluster, as well as45        direct access to cluster nodes for job scheduling. With HyperPod's lifecycle46        configuration support, you can customize the computing environment of the clusters to meet47        your specific requirements. Additionally, by leveraging the Amazon SageMaker AI distributed training48        libraries, you can optimize the clusters' performance on AWS computing and network49        resources. To learn more, see50Orchestrating SageMaker HyperPod clusters with Slurm51.52Amazon EKS support in SageMaker HyperPod53SageMaker HyperPod also integrates with Amazon EKS to enable large-scale training of foundation54        models on long-running and resilient compute clusters. This allows cluster admin users to55        provision HyperPod clusters and attach them to an EKS control plane, enabling56        dynamic capacity management, direct access to cluster instances, and resiliency57        capabilities. For data scientists, Amazon EKS support in HyperPod allows running58        containerized workloads for training foundation models, inference on the EKS cluster, and59        leveraging the job auto-resume capability for Kubeflow PyTorch training. The architecture60        involves a 1-to-1 mapping between an EKS cluster (control plane) and a HyperPod61        cluster (worker nodes) within a VPC, providing a tightly integrated solution for running62        large-scale ML workloads. To learn more, see63Orchestrating SageMaker HyperPod clusters with Amazon EKS64.65UltraServers with HyperPod66HyperPod with UltraServers delivers AI computing power by integrating67        NVIDIA superchips into a cohesive, high-performance infrastructure. Each NVL72 UltraServer 68        combines 18 instances with 72 NVIDIA Blackwell GPUs interconnected via NVLink, enabling69        faster inference and faster training performance compared to previous generation instances. This 70        architecture is particularly valuable for organizations working with trillion-parameter foundation 71        models, as the unified GPU memory allows entire models to remain within a single 72        NVLink domain, eliminating cross-node networking bottlenecks. HyperPod 73        enhances this hardware advantage74        with intelligent topology-aware scheduling that optimizes workload placement, automatic instance 75        replacement to minimize disruptions, and flexible deployment options that support both dedicated and 76        shared resource configurations. For teams pushing the boundaries of model size and performance, this77        integration provides the computational foundation needed to train and deploy the most advanced AI 78        models with unprecedented efficiency.79SageMaker HyperPod automatically optimizes instance placement across your UltraServers. 80        By default, HyperPod prioritizes all instances in one UltraServer before using a different one. 81        For example, if you want 14 instances and have 2 UltraServers in your plan, SageMaker AI uses all of the 82        instances in the first UltraServer. If you want 20 instances, SageMaker AI uses all 18 instances in the first83        UltraServer and then uses 2 more from the second.84AWS Regions supported by SageMaker HyperPod85SageMaker HyperPod is available in the following AWS Regions.86us-east-187us-east-288us-west-189us-west-290eu-central-191eu-north-192eu-west-193eu-west-294eu-south-295ap-south-196ap-southeast-197ap-southeast-298ap-southeast-399ap-southeast-4100ap-northeast-1101ap-northeast-2102sa-east-1103Topics104Amazon SageMaker HyperPod quickstart105Prerequisites for using SageMaker HyperPod106AWS Identity and Access Management for SageMaker HyperPod107Customer managed AWS KMS key encryption for SageMaker HyperPod108SageMaker HyperPod recipes109Orchestrating SageMaker HyperPod clusters with Slurm110Orchestrating SageMaker HyperPod clusters with Amazon EKS111Using topology-aware scheduling in Amazon SageMaker HyperPod112Deploying models on Amazon SageMaker HyperPod113HyperPod in Studio114SageMaker HyperPod references115Amazon SageMaker HyperPod release notes116Amazon SageMaker HyperPod AMI117Javascript is disabled or is unavailable in your browser.118To use the Amazon Web Services Documentation, Javascript must be enabled. Please refer to your browser's Help pages for instructions.119Document Conventions120Custom images121Quickstart122Did this page help you? - Yes123Thanks for letting us know we're doing a good job!124If you've got a moment, please tell us what we did right so we can do more of it.125Did this page help you? - No126Thanks for letting us know this page needs work. We're sorry we let you down.127If you've got a moment, please tell us how we can make the documentation better.