View a markdown version of this page

Ray Serve를 사용하여 JumpStart 모델 배포 - Amazon SageMaker AI

기계 번역으로 제공되는 번역입니다. 제공된 번역과 원본 영어의 내용이 상충하는 경우에는 영어 버전이 우선합니다.

Ray Serve를 사용하여 JumpStart 모델 배포

모델 가중치를 수동으로 다운로드하거나, 모델 로드 코드를 작성하거나, 서빙 컨테이너를 구성하지 않고도 JumpStart 모델을 HyperPod의 Ray Serve에 배포할 수 있습니다. PyPI 웹 사이트의 toolkit-for-ray-on-sagemaker-ai 라이브러리는 Ray Serve의 LLM API와 JumpStartModelLoaderCallback 통합되어 배포 시 JumpStart에서 모델 아티팩트 다운로드를 처리하는를 제공합니다.

사전 조건

IAM 권한 구성

는 SageMaker AI API를 JumpStartModelLoaderCallback 호출하여 모델 아티팩트 다운로드를 위해 미리 서명된 URLs을 검색합니다. Ray Serve 포드에는 다음 권한이 있는 IAM 역할이 필요합니다.

{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "sagemaker:CreateHubContentPresignedUrls", "Resource": "*" } ] }

Kubernetes 서비스 계정을이 IAM 역할에 매핑하는 Pod Identity 연결을 생성합니다. 자세한 내용은 Amazon EKS 사용 설명서의 EKS Pod Identity를 사용하여 IAM 역할을 수임하도록 Kubernetes 서비스 계정 구성을 참조하세요.

aws eks create-pod-identity-association \ --cluster-name eks-cluster-name \ --namespace namespace \ --service-account jumpstart-ray-serve \ --role-arn arn:aws:iam::account-id:role/role-name

serviceAccountName 필드를 사용하여 RayService 매니페스트에서이 서비스 계정을 참조합니다(다음 예제 참조).

모델 배포

다음 예시에서는와 LLMConfig 함께 Ray Serve의를 사용하여 JumpStart 모델을 배포합니다JumpStartModelLoaderCallback. 콜백은 서비스 복제본이 시작될 때 미리 서명된 URLs 사용하여 JumpStart에서 모델 아티팩트를 다운로드하고 이를 서비스 엔진(vLLM)에 로드하여 OpenAI 호환 API를 노출합니다.

from ray.serve.llm import LLMConfig, build_openai_app from ray.llm._internal.common.callbacks.base import CallbackConfig from toolkit_for_ray_on_sagemaker_ai import JumpStartModelLoaderCallback llm_config = LLMConfig( model_loading_config={ "model_id": "my-llm", "model_source": "placeholder", }, callback_config=CallbackConfig( callback_class=JumpStartModelLoaderCallback, callback_kwargs={ "jumpstart_model_id": "huggingface-reasoning-qwen3-4b", "region": "us-east-1", "accept_eula": False, # Set to True after reviewing the model's EULA }, ), accelerator_type="A10G", ) app = build_openai_app({"llm_configs": [llm_config]})

RayService 매니페스트의 import_path app에서 참조한 다음에 설명된 대로 배포합니다Ray Serve를 사용하여 모델 배포. headGroupSpec 및 serviceAccountName를 모두 이전 단계에서 생성된 IAM 역할과 연결된 서비스 계정으로 workerGroupSpecs 설정해야 합니다.

참고

모델 라이선스 및 EULA 요구 사항에 대한 자세한 내용은 파운데이션 모델 선택을 참조하세요.

엔드포인트에 도달

배포는에 대해 생성된 KubeRay 서비스를 통해 포트 8000에서 OpenAI 호환 채팅 완료 API를 노출합니다RayService. 빠른 테스트를 위해를 사용합니다kubectl port-forward.

kubectl port-forward svc/ray-service-head-svc 8000:8000

그런 http://localhost:8000다음에 요청을 보냅니다.

curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "my-llm", "messages": [{"role": "user", "content": "What is Ray Serve?"}] }'