SageMaker / Client / describe_cluster_event
describe_cluster_event¶
- SageMaker.Client.describe_cluster_event(**kwargs)¶
Retrieves detailed information about a specific event for a given HyperPod cluster. This functionality is only supported when the
NodeProvisioningModeis set toContinuous.See also: AWS API Documentation
Request Syntax
response = client.describe_cluster_event( EventId='string', ClusterName='string' )
- Parameters:
EventId (string) –
[REQUIRED]
The unique identifier (UUID) of the event to describe. This ID can be obtained from the
ListClusterEventsoperation.ClusterName (string) –
[REQUIRED]
The name or Amazon Resource Name (ARN) of the HyperPod cluster associated with the event.
- Return type:
dict
- Returns:
Response Syntax
{ 'EventDetails': { 'EventId': 'string', 'ClusterArn': 'string', 'ClusterName': 'string', 'InstanceGroupName': 'string', 'InstanceId': 'string', 'ResourceType': 'Cluster'|'InstanceGroup'|'Instance', 'EventTime': datetime(2015, 1, 1), 'EventDetails': { 'EventMetadata': { 'Cluster': { 'FailureMessage': 'string', 'EksRoleAccessEntries': [ 'string', ], 'SlrAccessEntry': 'string' }, 'InstanceGroup': { 'FailureMessage': 'string', 'AvailabilityZoneId': 'string', 'CapacityReservation': { 'Arn': 'string', 'Type': 'ODCR'|'CRG' }, 'SubnetId': 'string', 'SecurityGroupIds': [ 'string', ], 'AmiOverride': 'string' }, 'InstanceGroupScaling': { 'InstanceCount': 123, 'TargetCount': 123, 'MinCount': 123, 'FailureMessage': 'string' }, 'Instance': { 'CustomerEni': 'string', 'AdditionalEnis': { 'EfaEnis': [ 'string', ] }, 'InstanceRequirementsEniConfigurations': [ { 'CustomerEni': 'string', 'AdditionalEnis': { 'EfaEnis': [ 'string', ] } }, ], 'CapacityReservation': { 'Arn': 'string', 'Type': 'ODCR'|'CRG' }, 'FailureMessage': 'string', 'LcsExecutionState': 'string', 'NodeLogicalId': 'string' }, 'DatabaseConfiguration': { 'RollbackStatus': 'NotApplicable'|'Reverted'|'RevertFailed', 'Advisory': 'string', 'FailureMessage': 'string' }, 'SlurmHealth': { 'Component': 'Slurmdbd', 'Status': 'Healthy'|'Unhealthy', 'Reason': 'DaemonDown'|'DaemonDisabled'|'DbUnreachable' } } }, 'Description': 'string', 'EventLevel': 'Info'|'Warn'|'Error' } }
Response Structure
(dict) –
EventDetails (dict) –
Detailed information about the requested cluster event, including event metadata for various resource types such as
Cluster,InstanceGroup,Instance, and their associated attributes.EventId (string) –
The unique identifier (UUID) of the event.
ClusterArn (string) –
The Amazon Resource Name (ARN) of the HyperPod cluster associated with the event.
ClusterName (string) –
The name of the HyperPod cluster associated with the event.
InstanceGroupName (string) –
The name of the instance group associated with the event, if applicable.
InstanceId (string) –
The EC2 instance ID associated with the event, if applicable.
ResourceType (string) –
The type of resource associated with the event. Valid values are
Cluster,InstanceGroup, orInstance.EventTime (datetime) –
The timestamp when the event occurred.
EventDetails (dict) –
Additional details about the event, including event-specific metadata.
EventMetadata (dict) –
Metadata specific to the event, which may include information about the cluster, instance group, or instance involved.
Note
This is a Tagged Union structure. Only one of the following top level keys will be set:
Cluster,InstanceGroup,InstanceGroupScaling,Instance,DatabaseConfiguration,SlurmHealth. If a client receives an unknown member it will setSDK_UNKNOWN_MEMBERas the top level key, which maps to the name or tag of the unknown member. The structure ofSDK_UNKNOWN_MEMBERis as follows:'SDK_UNKNOWN_MEMBER': {'name': 'UnknownMemberName'}
Cluster (dict) –
Metadata specific to cluster-level events.
FailureMessage (string) –
An error message describing why the cluster level operation (such as creating, updating, or deleting) failed.
EksRoleAccessEntries (list) –
A list of Amazon EKS IAM role ARNs associated with the cluster. This is created by HyperPod on your behalf and only applies for EKS orchestrated clusters.
(string) –
SlrAccessEntry (string) –
The Service-Linked Role (SLR) associated with the cluster. This is created by HyperPod on your behalf and only applies for EKS orchestrated clusters.
InstanceGroup (dict) –
Metadata specific to instance group-level events.
FailureMessage (string) –
An error message describing why the instance group level operation (such as creating, scaling, or deleting) failed.
AvailabilityZoneId (string) –
The ID of the Availability Zone where the instance group is located.
CapacityReservation (dict) –
Information about the Capacity Reservation used by the instance group.
Arn (string) –
The Amazon Resource Name (ARN) of the Capacity Reservation.
Type (string) –
The type of Capacity Reservation. Valid values are
ODCR(On-Demand Capacity Reservation) orCRG(Capacity Reservation Group).
SubnetId (string) –
The ID of the subnet where the instance group is located.
SecurityGroupIds (list) –
A list of security group IDs associated with the instance group.
(string) –
AmiOverride (string) –
If you use a custom Amazon Machine Image (AMI) for the instance group, this field shows the ID of the custom AMI.
InstanceGroupScaling (dict) –
Metadata related to instance group scaling events.
InstanceCount (integer) –
The current number of instances in the group.
TargetCount (integer) –
The desired number of instances for the group after scaling.
MinCount (integer) –
Minimum instance count of the instance group.
FailureMessage (string) –
An error message describing why the scaling operation failed, if applicable.
Instance (dict) –
Metadata specific to instance-level events.
CustomerEni (string) –
The ID of the customer-managed Elastic Network Interface (ENI) associated with the instance.
AdditionalEnis (dict) –
Information about additional Elastic Network Interfaces (ENIs) associated with the instance.
EfaEnis (list) –
A list of Elastic Fabric Adapter (EFA) ENIs associated with the instance.
(string) –
InstanceRequirementsEniConfigurations (list) –
The ENI configurations for the instance types in the instance requirements, grouped by network interface category (for example, ENI-only or EFA with ENIs). At most one configuration per category.
(dict) –
The customer ENI and additional ENIs associated with a network interface category.
CustomerEni (string) –
The ID of the customer-managed Elastic Network Interface (ENI) associated with the instance type category.
AdditionalEnis (dict) –
Information about additional Elastic Network Interfaces (ENIs) associated with the instance type category.
EfaEnis (list) –
A list of Elastic Fabric Adapter (EFA) ENIs associated with the instance.
(string) –
CapacityReservation (dict) –
Information about the Capacity Reservation used by the instance.
Arn (string) –
The Amazon Resource Name (ARN) of the Capacity Reservation.
Type (string) –
The type of Capacity Reservation. Valid values are
ODCR(On-Demand Capacity Reservation) orCRG(Capacity Reservation Group).
FailureMessage (string) –
An error message describing why the instance creation or update failed, if applicable.
LcsExecutionState (string) –
The execution state of the Lifecycle Script (LCS) for the instance.
NodeLogicalId (string) –
The unique logical identifier of the node within the cluster. The ID used here is the same object as in the
BatchAddClusterNodesAPI.
DatabaseConfiguration (dict) –
Metadata specific to events about the external Slurm accounting database of the cluster.
RollbackStatus (string) –
Whether HyperPod restored the previous accounting database configuration after the change failed. Valid values:
NotApplicable: The change failed before HyperPod modified the cluster, for example because the database could not be reached or rejected the credentials, so there was nothing to restore.Reverted: The change failed after it was applied, and HyperPod restored the previous configuration. The cluster continues to use the previous accounting database.RevertFailed: The change failed and HyperPod could not restore the previous configuration, so Slurm accounting on the cluster might not be working.
This field is omitted when the change succeeds.
Advisory (string) –
Additional information about a change that succeeded, such as an action to take on the cluster.
FailureMessage (string) –
An error message describing why the accounting database change failed, and how to resolve it.
SlurmHealth (dict) –
Metadata specific to events about the health of the Slurm components on the controller node of the cluster.
Component (string) –
The Slurm component that the health information describes. The valid value is
Slurmdbd, the Slurm accounting daemon.Status (string) –
The health of the component. Valid values are
HealthyandUnhealthy.Reason (string) –
The reason the component is unhealthy. Valid values:
DaemonDown: The daemon is not running, so job accounting records are not being written.DaemonDisabled: The daemon is running and its accounting database is responding, but the daemon is not enabled to start automatically. Job accounting stops the next time the controller node restarts.DbUnreachable: The daemon is running, but its accounting database did not respond. Job accounting records might not be written.
This field is omitted when the component is healthy.
Description (string) –
A human-readable description of the event.
EventLevel (string) –
The severity level of the event. Valid values are
Info,Warn, andError.
Exceptions
SageMaker.Client.exceptions.ResourceNotFound