Amazon SageMaker HyperPod gives you a long-lived cluster of GPU or Trainium instances that AWS keeps healthy for you. You get a normal scheduler, Slurm or Amazon EKS, and you log in and run sbatch or kubectl as you would on your own cluster. Underneath, AWS runs a health agent on every node, can stress-test nodes before they join, and swaps out failed instances. With the right flag, it can also restart your training job on the repaired set of nodes.

This matters because large training runs fail all the time. Across hundreds of GPUs, one bad HBM stack, a flapping NVLink or a degraded network card will stop a synchronous job every few days or more often. The cost is not the failure itself. It is the hours between the failure and the restart, while a person diagnoses the node, gets a replacement and resubmits the job. HyperPod automates that path.

This page covers the cluster model, a working cluster definition, the parts of the resilience system and what each one actually does, a correct auto-resume entrypoint, the EKS view, a worked failure-cost example, and the trade-offs against running the same cluster yourself. Product details were checked against the AWS documentation in October 2026. HyperPod changes often, so check the API reference before relying on any field.

The cluster model

A HyperPod cluster: what AWS manages and what you ownSageMaker HyperPod control planeCreateCluster, UpdateCluster, node recovery, health logs to CloudWatchOrchestratorSlurm controller or EKSInstance group: workersml.p5.48xlarge x NInstance group: loginsmall CPU instancesprovisionOn every nodeHMA agent + your lifecycle scriptsS3 bucketlifecycle scripts, datasetsShared file systeme.g. FSx for Lustre: code, checkpointsEFA fabricNCCL traffic between workersOnCreateFault path: HMA flags node -> node drained -> replaced (or rebooted) -> job step re-run from last checkpoint
The parts of a HyperPod cluster. AWS manages the control plane and node replacement; you own the scheduler configuration, lifecycle scripts, storage layout and checkpoints.

A cluster is a set of instance groups. Each group has a name, an instance type such as ml.p5.48xlarge, a count, an IAM execution role and a lifecycle configuration. A typical Slurm cluster has three groups: a controller group with one CPU instance running slurmctld, a login group, and one or more worker groups with the accelerators. On EKS there is no controller group, because the EKS control plane plays that role and HyperPod nodes join your existing EKS cluster as worker nodes.

The instances are SageMaker-managed, so they do not appear in your EC2 console as ordinary instances. You reach them through AWS Systems Manager. The session target has the form sagemaker-cluster:<cluster-id>_<group-name>-<instance-id>, where the cluster ID is the last part of the cluster ARN, not the cluster name.

Lifecycle scripts are what turn bare instances into your cluster. You upload a directory of scripts to S3 and name an entry point, OnCreate, per instance group. HyperPod runs it on every node when the node is first created and again on every replacement node. That second case is the point. Anything a node needs, including Slurm configuration, mounts, users, Docker and drivers beyond the AMI, must be installed by these scripts and not by hand. AWS publishes sample scripts in its awsome-distributed-training repository on GitHub, and most teams start from those.

Defining a cluster

Clusters are created through the CreateCluster API. The fields below are the ones that change behaviour. Everything else can usually be left at its default.

{
  "ClusterName": "llm-train",
  "NodeRecovery": "Automatic",
  "VpcConfig": { "SecurityGroupIds": ["sg-0abc..."], "Subnets": ["subnet-0abc..."] },
  "InstanceGroups": [
    { "InstanceGroupName": "controller", "InstanceType": "ml.m5.2xlarge",
      "InstanceCount": 1, "ExecutionRole": "arn:aws:iam::111122223333:role/hp-exec",
      "LifeCycleConfig": { "SourceS3Uri": "s3://my-bucket/lifecycle/", "OnCreate": "on_create.sh" } },
    { "InstanceGroupName": "login", "InstanceType": "ml.m5.4xlarge",
      "InstanceCount": 1, "ExecutionRole": "arn:aws:iam::111122223333:role/hp-exec",
      "LifeCycleConfig": { "SourceS3Uri": "s3://my-bucket/lifecycle/", "OnCreate": "on_create.sh" } },
    { "InstanceGroupName": "workers", "InstanceType": "ml.p5.48xlarge",
      "InstanceCount": 16, "ThreadsPerCore": 1,
      "ExecutionRole": "arn:aws:iam::111122223333:role/hp-exec",
      "LifeCycleConfig": { "SourceS3Uri": "s3://my-bucket/lifecycle/", "OnCreate": "on_create.sh" },
      "OnStartDeepHealthChecks": ["InstanceStress", "InstanceConnectivity"] }
  ]
}
aws sagemaker create-cluster --cli-input-json file://cluster.json
aws sagemaker describe-cluster --cluster-name llm-train      # wait for InService
aws sagemaker list-cluster-nodes --cluster-name llm-train

Notes on the fields. For EKS you also pass Orchestrator with your EKS cluster ARN.

  • NodeRecovery is Automatic or None. With Automatic, HyperPod reboots or replaces faulty nodes itself. With None, it only labels them. AWS recommends Automatic and it is the default.
  • OnStartDeepHealthChecks accepts InstanceStress and InstanceConnectivity. They run when the group is created or updated, before nodes take work. They add time to cluster creation, but they catch bad hardware before it costs you a training run.
  • ThreadsPerCore set to 1 turns off hyperthreading.
  • TrainingPlanArn on a group attaches capacity reserved through SageMaker training plans. Without reserved capacity, a large P5 group can fail to provision, and replacement nodes can be slow to arrive in a busy Availability Zone.
  • Put all workers in one subnet, and so one Availability Zone. EFA traffic does not cross zones, and your security group must allow all traffic between members of the group, which EFA requires. AWS EFA, in depth covers the security-group rule and how to check that NCCL is really using EFA.

The resilience stack

Resilience is built from four separate parts. They are easy to confuse, and they cover different failures.

PartWhen it runsWhat it checks or does
Health-monitoring agent (HMA)continuously, on every GPU or Trainium nodeDCGM policy violations, errors in nvidia-smi output, EC2 platform error logs, GPU count (8 expected on ml.p5.48xlarge); Neuron monitor and device count on Trainium
Deep health checksat group create or update; on demand via StartClusterHealthCheck (documented for EKS)instance level: GPU and NVLink counts, DCGM diagnostics level 4, EFA latency and bandwidth; cluster level: NCCL tests across nodes
Node recoverywhen a check flags a node and NodeRecovery is Automaticreboots or replaces the instance, reruns lifecycle scripts, keeps the hostname
Auto-resumewhen a job step launched with the flag failswaits for replacement, then reruns the same srun step on the repaired allocation

The HMA is passive. It notices faults that are announced, such as an Xid error in the kernel log or a GPU dropping off the PCIe bus. It cannot see a GPU that is slow but reports nothing, or numerical corruption. Deep health checks are active. They load the hardware and measure it, and a node whose NCCL bandwidth is below threshold is marked faulty. Neither replaces the checks inside your training loop. Loss spikes, NaNs and stalled collectives are your job to detect. LLM training health checks, in depth covers the in-loop side, and GPU hardware faults explains what the Xid codes the agent reacts to mean.

The agent's detections go to CloudWatch Logs under /aws/sagemaker/Clusters/, one log stream per node. Deep health check results go to the cluster's log group under DeepHealthCheckResults/, and also to /var/log/aws/clusters/sagemaker-deep-health-check.log on the node. Set up a CloudWatch alarm on detection events. Every replacement deserves a look, because a node slot that keeps failing points at a systemic issue such as cooling, firmware or a job that triggers a driver bug.

Auto-resume done right

On Slurm, auto-resume is opt-in for each job step. You add --auto-resume=1 to srun inside an exclusive allocation. If the step fails because of a hardware fault, HyperPod replaces the bad nodes and runs the same step again. Three details decide whether that works:

  • Only the srun step with the flag is retried, not other steps in the same allocation. All environment setup must happen inside that one step, because a replacement node starts clean.
  • Do not trust $SLURM_JOB_NODELIST after a resume. AWS documents that it can be out of date. Ask scontrol for the job's current node list and derive the rendezvous address from it each time the step starts.
  • Auto-resume restarts your program. It does not save or restore anything. The job resumes from whatever checkpoint your code last wrote, so checkpoint frequency sets how much work you lose.

Here is a clean entrypoint. AWS's own sample has comments after line-continuation backslashes, which breaks the continuation in bash, so this version puts each step on its own line:

#!/bin/bash
# train_entry.sh -- run as: srun --auto-resume=1 train_entry.sh
set -euo pipefail
source /fsx/envs/train/bin/activate          # shared venv: same on replacement nodes

NODE_LIST=$(scontrol show jobid="$SLURM_JOBID" | awk -F= '/ NodeList=/{print $2}')
MASTER_NODE=$(scontrol show hostnames "$NODE_LIST" | head -n 1)
MASTER_ADDR=$(scontrol show node="$MASTER_NODE" | awk -F= '/NodeAddr=/{print $2}' | awk '{print $1}')

exec torchrun --nnodes="$SLURM_JOB_NUM_NODES" --nproc_per_node=8 \
     --rdzv_backend=c10d --rdzv_endpoint="$MASTER_ADDR:29500" \
     --rdzv_id="$SLURM_JOBID" train.py --ckpt-dir /fsx/ckpt/run42
#!/bin/bash
#SBATCH --nodes=16
#SBATCH --exclusive
srun --auto-resume=1 /fsx/code/train_entry.sh

The training script must resume by itself: on start, find the newest complete checkpoint and load it. Write checkpoints atomically, to a temporary name that is renamed when complete, so a crash during a save never leaves a half-written latest checkpoint.

def latest_complete(ckpt_dir):
    done = [d for d in os.listdir(ckpt_dir) if d.startswith("step_") and
            os.path.exists(os.path.join(ckpt_dir, d, "COMPLETE"))]
    return max(done, key=lambda d: int(d.split("_")[1]), default=None)

start = latest_complete(args.ckpt_dir)
if start:
    load_checkpoint(model, optimizer, os.path.join(args.ckpt_dir, start))

One caveat: if your Slurm nodes have generic resources (GRES) configured, Slurm does not allow the allocation to change. In that case HyperPod requeues the job and it restarts from the beginning of the batch script. Your code then resumes from its checkpoint, but queue wait time is added. Distributed checkpointing, in depth covers sharded saves that keep checkpoint stalls short enough to checkpoint often.

Manual recovery and the EKS view

Automatic recovery reacts to faults it can see. When you know a node is bad and the agent does not, for example because NCCL bandwidth tests show it running slow, trigger recovery yourself. From the Slurm controller:

scontrol update node=ip-10-1-2-3 state=fail reason="Action:Replace"
scontrol update node=ip-10-1-2-3 state=fail reason="Action:Reboot"

Or through the API, which works for both orchestrators:

aws sagemaker batch-replace-cluster-nodes --cluster-name llm-train --node-ids i-0123456789abcdef0
aws sagemaker batch-reboot-cluster-nodes  --cluster-name llm-train --node-ids i-0123456789abcdef0

Reboot for software trouble: hung processes, a wedged driver, memory leaks. Replace for hardware trouble: repeated Xid errors, failed DCGM diagnostics, low bandwidth that persists after a reboot. While a replacement is in progress, do not change the node's state again or restart slurmctld. AWS warns that this can make the replacement fail.

On EKS, the same state appears as Kubernetes metadata. A faulty node gets the label sagemaker.amazonaws.com/node-health-status, fault labels sagemaker.amazonaws.com/fault-types and sagemaker.amazonaws.com/fault-reasons, and the taint sagemaker.amazonaws.com/node-health-status=Unschedulable:NoSchedule, so new pods avoid it. Watch for nodes in that state with:

kubectl get nodes -L sagemaker.amazonaws.com/node-health-status,sagemaker.amazonaws.com/fault-types

Worked example: what a failure costs

Here is what resilience is worth in numbers. Take a 128-node job on ml.p5.48xlarge, 1,024 H100 GPUs, for 30 days. Meta's Llama 3 report counted 419 unexpected interruptions in 54 days on 16,384 H100s; scaled down, that is about one failure every two days here. Fifteen failures in total. The figures for detection and replacement below are assumptions to show the method. Measure your own.

Phase per failureManual clusterHyperPod with auto-resume
Notice the job died1 to 8 h (night, weekend); assume 3 hminutes
Find and drain the bad node30 minautomatic
Get and set up a replacement1 hassume 30 min (capacity and lifecycle scripts)
Lost work since last checkpointaverage half the interval: 30 min at 1 h checkpointssame: 30 min
Total per failureabout 5 habout 1 h

Fifteen failures at 5 hours is 75 hours of the whole cluster sitting idle, about 10% of the month. At 1 hour each it is 15 hours, about 2%. Two lessons follow. First, most of the gain comes from removing the human from the loop. Second, the remaining loss is dominated by checkpoint interval and replacement time. Halving the checkpoint interval to 30 minutes saves another 15 minutes per failure, which is worth doing if your checkpoints are fast. Reserved capacity matters too. If no spare P5 is free in the zone, replacement waits, and all 1,024 GPUs wait with it.

Failure modes

  • Hand-configured nodes. Someone installs a library by hand on the workers. The first replacement node lacks it, and the resumed job crashes again. Put everything in lifecycle scripts or a shared environment on the shared file system.
  • Lifecycle script failures. A script that fails, for example on a package mirror outage, makes node creation fail. Keep scripts idempotent, pin versions, avoid network fetches where you can, and log to a place you can read when the node never comes up.
  • Silent slowness. One slow GPU or link slows the whole synchronous job without triggering anything. Track step time per rank and run on-demand health checks when it drifts.
  • Checkpoints on local disk. A replaced node takes its NVMe contents with it. Write checkpoints to shared storage or S3.
  • Resume loops. A software bug, such as a bad batch that always produces a NaN, looks like a crash and can trigger repeated restarts from the same checkpoint. Cap restarts in your code and alert on repeats.

Trade-offs

Against a cluster you build yourself on EC2, with AWS ParallelCluster or self-managed EKS, HyperPod adds the health agent, deep checks, automatic replacement and auto-resume. It costs more per instance-hour than the same EC2 capacity, and you accept HyperPod AMIs, or custom AMIs that meet its rules, and its instance-type list. If you already have strong automation for node health and replacement, the premium buys less. If you do not, it is usually cheaper than building it.

Against SageMaker training jobs, which start a fresh cluster per job, HyperPod is persistent. You keep nodes warm between runs, debug interactively and run many jobs on one pool. The cost is that you pay for idle nodes. Use training jobs for occasional, self-contained runs, and HyperPod for a team that trains continuously.

Slurm or EKS? Pick the one your team already runs. Slurm is simpler for a handful of researchers sharing a pool. EKS fits a platform team that also serves models and wants Kubernetes tooling everywhere.

What to do next

  1. Start from AWS's sample lifecycle scripts, put them in version control, and change them only through commits.
  2. Create a small cluster with NodeRecovery set to Automatic and both deep health checks enabled. Read the deep health check log on one node.
  3. Write the entrypoint and resume logic above, then test it: mark a node with reason="Action:Replace" during a run and confirm the job resumes from its checkpoint.
  4. Measure your checkpoint write time and set the interval from it.
  5. Alert on health agent detection events and track per-rank step time.
  6. Reserve capacity, for example with a training plan, before scaling the worker group up.
Key takeaway: HyperPod is a persistent Slurm or EKS cluster whose nodes AWS monitors, stress tests and replaces. Auto-resume reruns your job step on repaired nodes, but it only restarts from your last checkpoint. Put all node setup in lifecycle scripts, checkpoint to shared storage, and reserve capacity so replacements arrive quickly.