Migrating to API-driven Slurm configuration
If you have an existing SageMaker HyperPod Slurm cluster that uses the legacy
provisioning_parameters.json file for Slurm configuration, you can migrate
to the API-driven configuration model. With API-driven configuration, you define Slurm node
types, partition assignments, and Amazon FSx mounting directly in the CreateCluster and UpdateCluster API payloads, removing the need to manage a separate configuration
file in Amazon S3.
Topics
Before you begin
Before starting the migration, make sure the following prerequisites are met:
-
Running jobs. If you choose the default Slurm configuration strategy,
Merge, running jobs are not affected by the migration. If you chooseManagedorOverwrite, drain all Slurm partitions and wait for active jobs to complete. You can check the job queue by runningsqueueon the controller node. With any strategy, do not submit new jobs while the update is in progress. Keep the defaultMergestrategy for this migration. For more information, see Choose a Slurm configuration strategy. -
AMI compatibility. API-driven Slurm configuration requires a SageMaker HyperPod Amazon Machine Image (AMI) released in January 2026 or later. If your cluster already runs a qualifying AMI, you can skip Step 2: Upgrade the cluster software. Otherwise, upgrade the cluster software before proceeding.
-
Custom VPC (if using Amazon FSx). If your cluster mounts Amazon FSx for Lustre or Amazon FSx for OpenZFS filesystems, the cluster must use a custom VPC (specified through
VpcConfigin the API payload). The platform-managed VPC cannot reach customer Amazon FSx resources. -
IAM permissions. The IAM role used to call the SageMaker AI API must have permissions for
sagemaker:UpdateClusterandsagemaker:DescribeCluster. The cluster execution role must have the managed AmazonSageMakerClusterInstanceRolePolicy attached. -
Single-controller clusters only. This migration guide applies to clusters with a single controller (head) node. If your cluster uses multiple controller nodes (multi-head configuration), see Clusters with multiple controller nodes (multi-head) before proceeding.
Clusters with multiple controller nodes (multi-head)
If your cluster is configured with multiple controller (head) nodes using the SageMaker HyperPod multi-head node support, you cannot fully migrate to the API-driven configuration model at this time.
The multi-head node architecture relies on the slurm_configurations block
in provisioning_parameters.json to configure several components that have no
equivalent in the API-driven approach:
"slurm_configurations": { "slurm_database_secret_arn": "$SLURM_DB_SECRET_ARN", "slurm_database_endpoint": "$SLURM_DB_ENDPOINT_ADDRESS", "slurm_shared_directory": "/fsx", "slurm_database_user": "$DB_USER_NAME", "slurm_sns_arn": "$SLURM_SNS_FAILOVER_TOPIC_ARN" }
The following table lists the fields in slurm_configurations and their
role in the multi-head node architecture.
| Field | Purpose | API equivalent |
|---|---|---|
slurm_database_secret_arn |
AWS Secrets Manager secret ARN for the external Slurm accounting database credentials | None |
slurm_database_endpoint |
Amazon RDS for MariaDB endpoint for Slurm accounting data (job records, metering) | None |
slurm_database_user |
Database user name for the Slurm accounting database | None |
slurm_shared_directory |
Shared Amazon FSx for Lustre directory used by multiple controller nodes to replicate Slurm state and configuration | None |
slurm_sns_arn |
Amazon SNS topic ARN for controller failover notifications (Slurm controller
ON/OFF status changes) |
None |
These fields configure the external Slurm accounting database
(slurmdbd), the shared filesystem for controller state replication, and
the SNS-based failover notification mechanism. Together, they enable the automatic
failover behavior between primary and backup controller nodes.
Because the API-driven SlurmConfig and Orchestrator.Slurm
schema does not include parameters for these components, migrating a multi-head cluster
to the API-driven approach would result in the loss of:
-
External Slurm accounting database connectivity (job history, metering data)
-
Automatic controller failover between primary and backup head nodes
-
SNS notifications for controller status changes
What you can do
If you have a multi-head node cluster and want to adopt parts of the API-driven approach, you have the following options:
-
Continue using
provisioning_parameters.json. This is the recommended approach for multi-head clusters. The legacy configuration remains fully supported and there is no forced migration. Your cluster continues to operate as expected. -
Migrate to a single-controller cluster. If you no longer require multi-head node high availability and are willing to give up the external accounting database and controller failover, you can restructure your cluster to use a single controller node and then follow the migration steps in this guide. This is a significant architectural change and should be evaluated carefully.
Important
Do not remove provisioning_parameters.json from Amazon S3 on a
multi-head node cluster. Doing so will break the external database configuration
and controller failover mechanism, which can lead to cluster instability.
Note
Support for multi-head node configuration parameters in the API-driven approach may be added in a future release. Check the Amazon SageMaker HyperPod release notes for updates.
What changes during migration
The following table summarizes what changes when you migrate from the legacy approach to API-driven configuration.
| Aspect | Before migration | After migration |
|---|---|---|
| Slurm node types | Defined in provisioning_parameters.json |
Defined in SlurmConfig within each instance group |
| Partition mapping | Defined in provisioning_parameters.json |
Defined in SlurmConfig.PartitionNames |
| Amazon FSx configuration | Defined in provisioning_parameters.json |
Defined in InstanceStorageConfigs per instance group |
| Configuration storage | Customer-managed Amazon S3 bucket | Managed by HyperPod |
| Update mechanism | Edit the Amazon S3 file, then call UpdateCluster |
Single UpdateCluster API call, no Amazon S3 file |
| Configuration strategy | Not applicable | Controlled by Orchestrator.Slurm.SlurmConfigStrategy (default
Merge) |
Note
Migration does not change your cluster's instance types, instance counts, lifecycle scripts, or execution roles. Only the source of Slurm topology configuration changes.
Note
If you plan to later migrate the cluster to continuous provisioning
(NodeProvisioningMode set to Continuous), keep the default
Merge strategy. Clusters using continuous provisioning support only
the Merge behavior and reject UpdateCluster requests that
set SlurmConfigStrategy explicitly.
Step 1: Back up your current configuration
Before making any changes, create backups of your existing configuration.
Download provisioning_parameters.json from Amazon S3:
aws s3 cp s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src/provisioning_parameters.json ./backup/
Note
If your instance groups use different SourceS3Uri locations, back up
provisioning_parameters.json from each of them. In Step 4, you
remove the file from every SourceS3Uri location, so make sure you have a
backup of each copy before you proceed.
Export the current slurm.conf from the controller node:
aws ssm start-session --targetmi-controller-instance-idsudo cp /opt/slurm/etc/slurm.conf ~/slurm.conf.backup
Record your current instance groups:
aws sagemaker describe-cluster \ --cluster-namemy-hyperpod-cluster\ --query "InstanceGroups[*].{Name:InstanceGroupName,Type:InstanceType,Count:CurrentCount}"
Keep these backups available throughout the migration process. You need them if you need to recover from a failed migration.
Step 2: Upgrade the cluster software
API-driven Slurm configuration requires a SageMaker HyperPod AMI released in January 2026 or later. If your cluster already runs a qualifying AMI, skip this step. Otherwise, upgrade the cluster software first.
Important
Before you upgrade the cluster software, update the lifecycle scripts in your Amazon S3
bucket to the latest version of the HyperPod sample lifecycle scriptsUpdateClusterSoftware re-runs your lifecycle
scripts on the new AMI, and scripts based on older versions of the samples can fail
on newer AMIs, which causes the upgrade to fail. If your scripts are already based on
a recent version, no change is needed.
Important
UpdateClusterSoftware replaces instance root volumes with the updated
AMI. Back up any data stored on root volumes to Amazon S3 or Amazon FSx before running this
command. For more information, see Update the SageMaker HyperPod platform software of a cluster.
aws sagemaker update-cluster-software \ --cluster-namemy-hyperpod-cluster
Wait for the update to complete:
aws sagemaker describe-cluster \ --cluster-namemy-hyperpod-cluster\ --query "ClusterStatus"
Proceed to the next step when the cluster status returns to
InService.
Step 3: Map your legacy configuration to the API format
Convert the fields in your provisioning_parameters.json file to the
corresponding API parameters.
Field mapping reference
The following table maps each legacy provisioning_parameters.json
field to its corresponding API parameter.
Legacy field (provisioning_parameters.json) |
API parameter |
|---|---|
controller_group |
Instance group with SlurmConfig.NodeType set to
"Controller" |
login_group |
Instance group with SlurmConfig.NodeType set to
"Login" |
worker_groups[].instance_group_name |
Instance group with SlurmConfig.NodeType set to
"Compute" |
worker_groups[].partition_name |
SlurmConfig.PartitionNames |
fsx_dns_name |
InstanceStorageConfigs[].FsxLustreConfig.DnsName |
fsx_mountname |
InstanceStorageConfigs[].FsxLustreConfig.MountName |
No legacy field (mount directory, typically /fsx) |
InstanceStorageConfigs[].FsxLustreConfig.MountPath |
Note
The slurm_configurations block used in multi-head node clusters
(containing slurm_database_secret_arn,
slurm_database_endpoint, slurm_database_user,
slurm_shared_directory, and slurm_sns_arn) has no
equivalent in the API-driven approach. If your
provisioning_parameters.json includes this block, see Clusters with multiple controller nodes (multi-head) before
proceeding.
Example conversion
The following example shows a legacy provisioning_parameters.json file
and the equivalent API-driven UpdateCluster request payload.
Legacy provisioning_parameters.json:
{ "version": "1.0.0", "workload_manager": "slurm", "controller_group": "controller-machine", "login_group": "login-group", "worker_groups": [ { "instance_group_name": "gpu-compute", "partition_name": "gpu-training" }, { "instance_group_name": "cpu-compute", "partition_name": "cpu-batch" } ], "fsx_dns_name": "fs-0abc123def456789.fsx.us-west-2.amazonaws.com", "fsx_mountname": "abcdefgh" }
Equivalent API-driven UpdateCluster request
(update_cluster_migration.json):
{ "ClusterName": "my-hyperpod-cluster", "InstanceGroups": [ { "InstanceGroupName": "controller-machine", "InstanceType": "ml.c5.xlarge", "InstanceCount": 1, "SlurmConfig": { "NodeType": "Controller" }, "LifeCycleConfig": { "SourceS3Uri": "s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src", "OnCreate": "on_create.sh" }, "ExecutionRole": "arn:aws:iam::111122223333:role/HyperPodExecutionRole", "InstanceStorageConfigs": [ { "EbsVolumeConfig": { "VolumeSizeInGB": 500 } } ] }, { "InstanceGroupName": "login-group", "InstanceType": "ml.m5.xlarge", "InstanceCount": 1, "SlurmConfig": { "NodeType": "Login" }, "LifeCycleConfig": { "SourceS3Uri": "s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src", "OnCreate": "on_create.sh" }, "ExecutionRole": "arn:aws:iam::111122223333:role/HyperPodExecutionRole" }, { "InstanceGroupName": "gpu-compute", "InstanceType": "ml.g5.12xlarge", "InstanceCount": 4, "SlurmConfig": { "NodeType": "Compute", "PartitionNames": ["gpu-training"] }, "InstanceStorageConfigs": [ { "FsxLustreConfig": { "DnsName": "fs-0abc123def456789.fsx.us-west-2.amazonaws.com", "MountPath": "/fsx", "MountName": "abcdefgh" } } ], "LifeCycleConfig": { "SourceS3Uri": "s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src", "OnCreate": "on_create.sh" }, "ExecutionRole": "arn:aws:iam::111122223333:role/HyperPodExecutionRole" }, { "InstanceGroupName": "cpu-compute", "InstanceType": "ml.c5.4xlarge", "InstanceCount": 2, "SlurmConfig": { "NodeType": "Compute", "PartitionNames": ["cpu-batch"] }, "InstanceStorageConfigs": [ { "FsxLustreConfig": { "DnsName": "fs-0abc123def456789.fsx.us-west-2.amazonaws.com", "MountPath": "/fsx", "MountName": "abcdefgh" } } ], "LifeCycleConfig": { "SourceS3Uri": "s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src", "OnCreate": "on_create.sh" }, "ExecutionRole": "arn:aws:iam::111122223333:role/HyperPodExecutionRole" } ], "Orchestrator": { "Slurm": { "SlurmConfigStrategy": "Merge" } } }
Note the following about this conversion:
-
Each instance group includes a
SlurmConfigblock with the appropriateNodeType(Controller,Login, orCompute). -
Partition names from
worker_groups[].partition_namemove toSlurmConfig.PartitionNameson the corresponding compute instance group. -
Amazon FSx configuration moves from the top-level
fsx_dns_nameandfsx_mountnamefields toInstanceStorageConfigs.FsxLustreConfigon each instance group that needs the filesystem mounted. With the API-driven approach, you can configure Amazon FSx per instance group rather than cluster-wide. -
SlurmConfigStrategyis set toMerge, which is the default. This preserves any manual edits you have made toslurm.confon the controller node and is required if you later migrate the cluster to continuous provisioning. For more information about configuration strategies, see Choose a Slurm configuration strategy.
Important
Make sure the InstanceGroupName, InstanceType, and
ExecutionRole values match your existing cluster configuration.
You can retrieve these values by running aws sagemaker describe-cluster
--cluster-name my-hyperpod-cluster.
Step 4: Remove provisioning_parameters.json from Amazon S3
Removing the provisioning_parameters.json file from Amazon S3 signals
HyperPod to use the API-driven configuration instead of the legacy file-based
approach.
aws s3 rm s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src/provisioning_parameters.json
Verify the file has been removed:
aws s3 ls s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src/
If your instance groups use different SourceS3Uri locations, remove the
file from each of them.
Important
HyperPod does not allow a cluster to use both
provisioning_parameters.json and the API-driven
SlurmConfig. Remove the file from Amazon S3 before you apply the
API-driven configuration in Step 5: Apply the API-driven configuration. If the file is
still present when you call UpdateCluster, the update fails and rolls
back.
Important
Make sure you have a local backup of provisioning_parameters.json
before deleting it. You need this file if you need to recover from a failed
migration.
Step 5: Apply the API-driven configuration
Run the UpdateCluster API with the JSON request file you prepared in Step 3: Map your legacy configuration to the API format:
aws sagemaker update-cluster \ --cli-input-jsonfile://update_cluster_migration.json
Monitor the update progress:
aws sagemaker describe-cluster \ --cluster-namemy-hyperpod-cluster\ --query "{Status: ClusterStatus, Message: FailureMessage}"
Wait for the cluster status to return to InService before proceeding. The
update typically takes several minutes depending on the number of instance groups and
nodes in your cluster.
Step 6: Verify the migration
After the cluster status returns to InService, verify that the Slurm
configuration was applied correctly.
Check Slurm partitions. Connect to the controller node using SSM Session Manager and verify the partition layout:
aws ssm start-session --targetmi-controller-instance-id
sinfo
Expected output:
PARTITION AVAIL TIMELIMIT NODES STATE NODELIST dev* up infinite 6 idle gpu-compute-[1-4],cpu-compute-[1-2] gpu-training up infinite 4 idle gpu-compute-[1-4] cpu-batch up infinite 2 idle cpu-compute-[1-2]
Check node-to-partition assignments.
scontrol show nodes | grep -E "NodeName|Partitions"
Check Amazon FSx mounts (if configured). On a compute node, verify the filesystem is mounted:
df -h | grep fsx
Expected output:
fs-0abc123def456789.fsx.us-west-2.amazonaws.com@tcp:/abcdefgh 1.2T 12G 1.2T 1% /fsx
Submit a test job.
sbatch --partition=gpu-training --wrap="hostname && nvidia-smi"
Check the job output:
squeue cat slurm-*.out
Recovering from a failed migration
If the migration does not complete successfully, HyperPod automatically rolls
back the cluster to its previous state. The cluster returns to InService and
continues to use the legacy configuration.
To recover, restore provisioning_parameters.json to its original Amazon S3
location so that any new nodes provision with the legacy configuration:
aws s3 cp ./backup/provisioning_parameters.json \ s3://DOC-EXAMPLE-BUCKET/lifecycle-script-directory/src/provisioning_parameters.json
Then review the failure reason, correct your request payload, and retry from Step 4: Remove provisioning_parameters.json from Amazon S3:
aws sagemaker describe-cluster \ --cluster-namemy-hyperpod-cluster\ --query "{Status: ClusterStatus, Message: FailureMessage}"
Important
After a migration completes successfully, you cannot revert the cluster to the
legacy provisioning_parameters.json configuration. Omitting
SlurmConfig from a later UpdateCluster request does not
revert the cluster; the existing API-driven Slurm configuration is preserved. Verify
the migration (Step 6: Verify the migration)
before you delete your local backups.
Post-migration considerations
After a successful migration, consider the following updates to your environment.
Simplify lifecycle scripts
With API-driven configuration, HyperPod handles Slurm topology setup and Amazon FSx mounting automatically. You can remove the corresponding logic from your lifecycle scripts and keep only custom setup steps such as user creation, package installation, and environment configuration.
The following example shows a minimal on_create.sh lifecycle script
for clusters using API-driven configuration:
#!/bin/bash set -e echo "=== HyperPod Lifecycle Script ===" echo "Timestamp: $(date)" echo "Hostname: $(hostname)" # Custom setup only - Slurm configuration and FSx mounting # are handled by HyperPod based on the API payload. # Example: Install additional packages # sudo apt-get update && sudo apt-get install -y htop vim # Example: Create users # sudo useradd -m -s /bin/bash researcher # Example: Set environment variables # echo 'export NCCL_DEBUG=INFO' >> /etc/profile.d/nccl.sh echo "=== Lifecycle Script Completed ==="
Update automation workflows
If you have automation scripts or CI/CD pipelines that modify
provisioning_parameters.json in Amazon S3, update them to use the
UpdateCluster API instead. Remove any Amazon S3 file management logic
related to provisioning_parameters.json.
Choose a Slurm configuration strategy
After you migrate, you can choose a SlurmConfigStrategy that controls
how HyperPod manages the relationship between the API-declared Slurm topology
and the actual slurm.conf on the controller node. The strategy
determines who owns the Slurm topology—you or the API.
| Strategy | Source of truth | Behavior |
|---|---|---|
Merge (default) |
slurm.conf |
HyperPod adds API-declared partitions to slurm.conf
without removing manual edits. Use this if you tune
slurm.conf directly for advanced parameters the API does
not expose. Required if you plan to migrate to continuous provisioning. |
Managed |
API | HyperPod enforces the API-declared state and detects drift. If
someone edits slurm.conf outside the API, the next
UpdateCluster fails and rolls back, and
DescribeCluster reports the drift in
FailureMessage until it is resolved. |
Overwrite |
API | HyperPod enforces the API-declared state unconditionally,
overwriting any manual changes to slurm.conf. Use this for
recovery scenarios or strict API governance. |
You can change the strategy in a later UpdateCluster request by
including Orchestrator.Slurm.SlurmConfigStrategy. Clusters that use
continuous provisioning (NodeProvisioningMode set to
Continuous) support only the Merge behavior and reject
requests that set SlurmConfigStrategy. If you plan to migrate to
continuous provisioning, keep Merge. For more information, see Continuous provisioning for enhanced cluster operations with Slurm.
Tip
If you are unsure which strategy to use, start with Merge. It
matches the behavior of the legacy approach and preserves any manual
slurm.conf customizations. You can switch to Managed
or Overwrite later as your operational practices evolve, unless you
plan to migrate to continuous provisioning, which requires
Merge.
Troubleshooting
"Update required to use SlurmConfig in InstanceGroups"
Your cluster AMI does not include the updated agents required for API-driven configuration. Upgrade the cluster software:
aws sagemaker update-cluster-software \ --cluster-namemy-hyperpod-cluster
Wait for the cluster to return to InService, then retry the
migration.
"N Controller Groups found: SlurmConfig for InstanceGroups may only have one Controller Group."
Your API payload specifies more than one instance group with
SlurmConfig.NodeType set to "Controller". Make sure
exactly one instance group is the controller.
"Cluster X has no InstanceGroup with Controller node type."
Your API payload does not identify a controller instance group. Make sure exactly
one instance group has SlurmConfig.NodeType set to
"Controller".
"Partitions can only be assigned to Compute node types"
You specified PartitionNames on a controller or login instance group.
Remove PartitionNames from any instance group that is not a compute
node.
"The LifeCycleConfig cannot include both a SLURM Orchestrator Config and a provisioning_parameters.json file simultaneously."
The update workflow found provisioning_parameters.json in the
lifecycle script location while your request included SlurmConfig. The
update fails and rolls back, and DescribeCluster reports this message in
FailureMessage. Complete Step 4: Remove provisioning_parameters.json from Amazon S3, verify that no
copy of the file remains under the SourceS3Uri prefix of any instance
group, and retry UpdateCluster.
Slurm partitions not updated after migration
If sinfo does not show the expected partitions:
-
Check the Cluster Agent logs on the controller node:
sudo journalctl -u cluster-agent -
Verify that
SlurmConfig.PartitionNamesis specified for each compute instance group in your API payload. -
If using the
Managedstrategy, checkDescribeClusterFailureMessagefor a drift error. A drift error fails and rolls back the update until the conflict is resolved. You can switch toOverwritetemporarily to force the API state, then switch back toManaged.
Amazon FSx not mounted on nodes
Amazon FSx configuration changes only apply to new nodes. Existing nodes retain their original storage configuration. To apply Amazon FSx changes to all nodes in an instance group:
-
Scale the instance group down to 0 instances.
-
Scale the instance group back up with the new Amazon FSx configuration.
Also verify that your security groups allow the required ports:
-
Amazon FSx for Lustre: TCP ports 988, 1021–1023
-
Amazon FSx for OpenZFS: TCP port 2049
"Configuration drift detected."
Drift messages (for example, "Configuration drift detected.", "Partition
configuration drift detected.", or "Partition configuration mismatch detected.")
appear in the FailureMessage field of DescribeCluster after
an UpdateCluster fails and rolls back. They occur when using the
Managed strategy and HyperPod detects configuration drift—a
mismatch between the API-declared configuration and the actual
slurm.conf on the controller node. Drift occurs in the following
cases:
-
A partition exists in
slurm.confthat noSlurmConfigdefines. -
A node belongs to more than one partition.
-
A node is in a different partition than the API expects.
To resolve drift, use one of the following options:
-
Option 1: Switch to
Overwritestrategy temporarily to force the API-declared state, then switch back toManaged. -
Option 2: Switch to
Mergestrategy to preserve the manual edits. -
Option 3: Connect to the controller node, manually revert the changes in
slurm.confto match the expected state, runsudo scontrol reconfigure, and retry theUpdateClustercall withManaged.
Migration checklist
Use this checklist to track your migration progress:
-
Verify your cluster is not using multi-head node configuration.
-
If your cluster mounts Amazon FSx, confirm it uses a custom VPC.
-
Back up
provisioning_parameters.jsonfrom Amazon S3. -
Back up
slurm.conffrom the controller node. -
Document current instance groups and their configuration.
-
If using
ManagedorOverwrite, drain partitions and wait for jobs to complete (squeue). -
Do not submit new jobs while the update is in progress.
-
Update lifecycle scripts to the latest sample version before upgrading cluster software.
-
Upgrade cluster software if the AMI predates January 2026 (
UpdateClusterSoftware). -
Prepare the API-driven
UpdateClusterrequest payload. -
Remove
provisioning_parameters.jsonfrom every instance group'sSourceS3Urilocation and verify it is gone. -
Run
UpdateClusterwith the new payload. -
Verify Slurm partitions (
sinfo). -
Verify Amazon FSx mounts (
df -h). -
Submit test jobs to each partition.
-
Update lifecycle scripts to remove Slurm and Amazon FSx setup logic.
-
Update automation scripts and CI/CD pipelines.
-
Keep your
provisioning_parameters.jsonandslurm.confbackups until Step 6 verification passes.