Terjemahan disediakan oleh mesin penerjemah. Jika konten terjemahan yang diberikan bertentangan dengan versi bahasa Inggris aslinya, utamakan versi bahasa Inggris.
Menanggapi peristiwa batas waktu perubahan ukuran armada instans cluster Amazon EMR
Gambaran umum
Cluster Amazon EMR memancarkan peristiwa saat menjalankan operasi pengubahan ukuran misalnya cluster armada. Peristiwa batas waktu penyediaan dipancarkan saat Amazon EMR berhenti menyediakan Spot atau On-demand kapasitas untuk armada setelah batas waktu berakhir. Durasi batas waktu dapat dikonfigurasi oleh pengguna sebagai bagian dari spesifikasi perubahan ukuran untuk armada instans. Dalam skenario perubahan ukuran berturut-turut untuk armada instans yang sama, Amazon EMR memancarkan On-Demand provisioning
timeout - continuing resize peristiwa Spot
provisioning timeout - continuing resize or saat batas waktu untuk operasi pengubahan ukuran saat ini berakhir. Kemudian mulai menyediakan kapasitas untuk operasi pengubahan ukuran armada berikutnya.
Menanggapi peristiwa batas waktu perubahan ukuran armada instans
Sebaiknya tanggapi peristiwa batas waktu penyediaan dengan salah satu cara berikut:
-
Kunjungi kembali spesifikasi pengu bah ukuran dan coba lagi operasi pengubah ukuran. Karena kapasitas sering bergeser, cluster Anda akan berhasil mengubah ukuran segera setelah kapasitas Amazon EC2 tersedia. Kami menyarankan pelanggan untuk mengonfigurasi nilai yang lebih rendah untuk durasi batas waktu untuk pekerjaan yang memerlukan SLA yang lebih ketat.
-
Atau, Anda dapat:
-
Meluncurkan cluster baru dengan tipe instans yang beragam berdasarkan praktik terbaik misalnya dan fleksibilitas Zona Ketersediaan atau
-
Luncurkan cluster dengan On-demand kapasitas
-
-
Untuk peristiwa batas waktu penyediaan - melanjutkan perubahan ukuran, Anda juga dapat menunggu operasi pengubahan ukuran diproses. Amazon EMR akan terus memproses operasi pengubahan ukuran yang dipicu untuk armada secara berurutan, dengan menghormati spesifikasi pengubahan ukuran yang dikonfigurasi.
Anda juga dapat mengatur aturan atau tanggapan otomatis untuk acara ini seperti yang dijelaskan di bagian selanjutnya.
Pemulihan otomatis dari peristiwa batas waktu penyediaan
Anda dapat membangun otomatisasi sebagai respons terhadap peristiwa Amazon EMR dengan kode Spot
Provisioning timeout acara. Misalnya, AWS Lambda
fungsi berikut mematikan cluster EMR dengan armada instans yang menggunakan instans Spot untuk node Tugas, dan kemudian membuat cluster EMR baru dengan armada instans yang berisi jenis instans yang lebih beragam daripada permintaan asli. Dalam contoh ini, Spot Provisioning timeout peristiwa yang dipancarkan untuk node tugas akan memicu eksekusi fungsi Lambda.
contoh Contoh fungsi untuk merespons peristiwa batas waktu Spot Provisioning
// Lambda code with Python 3.10 and handler is lambda_function.lambda_handler // Note: related IAM role requires permission to use Amazon EMR import json import boto3 import datetime from datetime import timezone SPOT_PROVISIONING_TIMEOUT_EXCEPTION_DETAIL_TYPE = "EMR Instance Fleet Resize" SPOT_PROVISIONING_TIMEOUT_EXCEPTION_EVENT_CODE = ( "Spot Provisioning timeout" ) CLIENT = boto3.client("emr", region_name="us-east-1") # checks if the incoming event is 'EMR Instance Fleet Resize' with eventCode 'Spot provisioning timeout' def is_spot_provisioning_timeout_event(event): if not event["detail"]: return False else: return ( event["detail-type"] == SPOT_PROVISIONING_TIMEOUT_EXCEPTION_DETAIL_TYPE and event["detail"]["eventCode"] == SPOT_PROVISIONING_TIMEOUT_EXCEPTION_EVENT_CODE ) # checks if the cluster is eligible for termination def is_cluster_eligible_for_termination(event, describeClusterResponse): # instanceFleetType could be CORE, MASTER OR TASK instanceFleetType = event["detail"]["instanceFleetType"] # Check if instance fleet receiving Spot provisioning timeout event is TASK if (instanceFleetType == "TASK"): return True else: return False # create a new cluster by choosing different InstanceType. def create_cluster(event): # instanceFleetType cloud be CORE, MASTER OR TASK instanceFleetType = event["detail"]["instanceFleetType"] # the following two lines assumes that the customer that created the cluster already knows which instance types they use in original request instanceTypesFromOriginalRequestMaster = "m5.xlarge" instanceTypesFromOriginalRequestCore = "m5.xlarge" # select new instance types to include in the new createCluster request instanceTypesForTask = [ "m5.xlarge", "m5.2xlarge", "m5.4xlarge", "m5.8xlarge", "m5.12xlarge" ] print("Starting to create cluster...") instances = { "InstanceFleets": [ { "InstanceFleetType":"MASTER", "TargetOnDemandCapacity":1, "TargetSpotCapacity":0, "InstanceTypeConfigs":[ { 'InstanceType': instanceTypesFromOriginalRequestMaster, "WeightedCapacity":1, } ] }, { "InstanceFleetType":"CORE", "TargetOnDemandCapacity":1, "TargetSpotCapacity":0, "InstanceTypeConfigs":[ { 'InstanceType': instanceTypesFromOriginalRequestCore, "WeightedCapacity":1, } ] }, { "InstanceFleetType":"TASK", "TargetOnDemandCapacity":0, "TargetSpotCapacity":100, "LaunchSpecifications":{}, "InstanceTypeConfigs":[ { 'InstanceType': instanceTypesForTask[0], "WeightedCapacity":1, }, { 'InstanceType': instanceTypesForTask[1], "WeightedCapacity":2, }, { 'InstanceType': instanceTypesForTask[2], "WeightedCapacity":4, }, { 'InstanceType': instanceTypesForTask[3], "WeightedCapacity":8, }, { 'InstanceType': instanceTypesForTask[4], "WeightedCapacity":12, } ], "ResizeSpecifications": { "SpotResizeSpecification": { "TimeoutDurationMinutes": 30 } } } ] } response = CLIENT.run_job_flow( Name="Test Cluster", Instances=instances, VisibleToAllUsers=True, JobFlowRole="EMR_EC2_DefaultRole", ServiceRole="EMR_DefaultRole", ReleaseLabel="emr-6.10.0", ) return response["JobFlowId"] # terminated the cluster using clusterId received in an event def terminate_cluster(event): print("Trying to terminate cluster, clusterId: " + event["detail"]["clusterId"]) response = CLIENT.terminate_job_flows(JobFlowIds=[event["detail"]["clusterId"]]) print(f"Terminate cluster response: {response}") def describe_cluster(event): response = CLIENT.describe_cluster(ClusterId=event["detail"]["clusterId"]) return response def lambda_handler(event, context): if is_spot_provisioning_timeout_event(event): print( "Received spot provisioning timeout event for instanceFleet, clusterId: " + event["detail"]["clusterId"] ) describeClusterResponse = describe_cluster(event) shouldTerminateCluster = is_cluster_eligible_for_termination( event, describeClusterResponse ) if shouldTerminateCluster: terminate_cluster(event) clusterId = create_cluster(event) print("Created a new cluster, clusterId: " + clusterId) else: print( "Cluster is not eligible for termination, clusterId: " + event["detail"]["clusterId"] ) else: print("Received event is not spot provisioning timeout event, skipping")