

기계 번역으로 제공되는 번역입니다. 제공된 번역과 원본 영어의 내용이 상충하는 경우에는 영어 버전이 우선합니다.

# 엔드포인트 구성 생성
<a name="async-inference-create-endpoint-create-endpoint-config"></a>

모델을 만들었으면 [`CreateEndpointConfig`](https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_CreateEndpointConfig.html)를 사용하여 엔드포인트 구성을 생성하세요. Amazon SageMaker AI 호스팅 서비스는 이 구성을 사용하여 모델을 배포합니다. 구성에서, [`CreateModel`](https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_CreateModel.html)을 사용하여 만든 하나 이상의 모델을 식별하여 Amazon SageMaker AI에서 프로비저닝할 리소스를 배포할 수 있습니다. `AsyncInferenceConfig` 객체를 지정하고 `OutputConfig`에 대한 출력 Amazon S3 위치를 제공합니다. 예측 결과에 대한 알림을 전송할 [Amazon SNS](https://docs.aws.amazon.com/sns/latest/dg/welcome.html) 주제를 선택적으로 지정할 수 있습니다. Amazon SNS 주제에 대한 자세한 내용은 [Amazon SNS 구성](https://docs.aws.amazon.com/sns/latest/dg/sns-configuring.html)을 참조하세요.

엔드포인트 가용성을 개선하고 용량 부족 오류를 줄이기 위해 프로덕션 변형에서 *인스턴스 풀을* 구성할 수 있습니다. 인스턴스 풀을 사용하면 원하는 인스턴스 유형을 사용할 수 없는 경우 SageMaker AI가 우선 순위가 낮은 유형으로 자동 폴백되도록 정렬된 인스턴스 유형 목록을 지정할 수 있습니다. 자세한 내용은 [인스턴스 풀을 사용하여 여러 인스턴스 유형에 배포](realtime-endpoints-heterogeneous.md) 단원을 참조하십시오.

다음 예제는 AWS SDK for Python (Boto3)을 사용하여 엔드포인트 구성을 생성하는 방법을 보여줍니다.

```
import datetime
from time import gmtime, strftime

# Create an endpoint config name. Here we create one based on the date  
# so it we can search endpoints based on creation time.
endpoint_config_name = f"XGBoostEndpointConfig-{strftime('%Y-%m-%d-%H-%M-%S', gmtime())}"

# The name of the model that you want to host. This is the name that you specified when creating the model.
model_name={{'<The_name_of_your_model>'}}

create_endpoint_config_response = sagemaker_client.create_endpoint_config(
    EndpointConfigName=endpoint_config_name, # You will specify this name in a CreateEndpoint request.
    # List of ProductionVariant objects, one for each model that you want to host at this endpoint.
    ProductionVariants=[
        {
            "VariantName": {{"variant1"}}, # The name of the production variant.
            "ModelName": model_name, 
            "InstanceType": {{"ml.m5.xlarge"}}, # Specify the compute instance type.
            "InitialInstanceCount": {{1}} # Number of instances to launch initially.
        }
    ],
    AsyncInferenceConfig={
        "OutputConfig": {
            # Location to upload response outputs when no location is provided in the request.
            "S3OutputPath": f"s3://{s3_bucket}/{bucket_prefix}/output"
            # (Optional) specify Amazon SNS topics
            "NotificationConfig": {
                "SuccessTopic": "arn:aws:sns:{{aws-region:account-id:topic-name}}",
                "ErrorTopic": "arn:aws:sns:{{aws-region:account-id:topic-name}}",
            }
        },
        "ClientConfig": {
            # (Optional) Specify the max number of inflight invocations per instance
            # If no value is provided, Amazon SageMaker will choose an optimal value for you
            "MaxConcurrentInvocationsPerInstance": 4
        }
    }
)

print(f"Created EndpointConfig: {create_endpoint_config_response['EndpointConfigArn']}")
```

앞서 언급한 예시에서는 `AsyncInferenceConfig` 필드에 `OutputConfig`에 대해 다음 키를 지정합니다.
+ `S3OutputPath`: 요청에 위치가 제공되지 않은 경우 응답 출력을 업로드할 위치입니다.
+ `NotificationConfig`: (선택 사항) 추론 요청이 성공했을 경우(`SuccessTopic`) 또는 실패할 경우(`ErrorTopic`) 알림을 게시하는 SNS 주제.

`AsyncInferenceConfig` 필드에서 `ClientConfig`에 대한 다음과 같은 선택적 인수를 지정할 수도 있습니다.
+ `MaxConcurrentInvocationsPerInstance`: (선택 사항) SageMaker AI 클라이언트가 모델 컨테이너로 보낸 최대 동시 요청 수입니다.