View a markdown version of this page

Ray Serve を使用したモデルのデプロイ - Amazon SageMaker AI

翻訳は機械翻訳により提供されています。提供された翻訳内容と英語版の間で齟齬、不一致または矛盾がある場合、英語版が優先します。

Ray Serve を使用したモデルのデプロイ

RayService リソースを使用してモデルをデプロイします。serveConfigV2 フィールドは、Serve アプリケーションとそのデプロイを保持し、KubeRay はそれらを実行する Ray クラスターを作成します。HyperPod は KubeRay を変更しないため、すでに実行RayServiceしている は引き続き機能します。

RayService マニフェスト

次のマニフェストは、単一の GPU-backed デプロイで 1 つの Serve アプリケーションを実行します。を作業ディレクトリのモジュールとアプリケーションオブジェクトimport_pathに置き換えます。

apiVersion: ray.io/v1 kind: RayService metadata: name: my-service namespace: my-namespace spec: serveConfigV2: | applications: - name: my-app import_path: my_module:app route_prefix: / deployments: - name: Model num_replicas: 2 ray_actor_options: num_gpus: 1 rayClusterConfig: rayVersion: "2.56.1" headGroupSpec: rayStartParams: dashboard-host: "0.0.0.0" template: spec: containers: - name: ray-head image: rayproject/ray:2.56.1 ports: - { containerPort: 8265, name: dashboard } - { containerPort: 8000, name: serve } workerGroupSpecs: - groupName: gpu-workers replicas: 1 rayStartParams: {} template: spec: nodeSelector: node.kubernetes.io/instance-type: ml.g5.xlarge containers: - name: ray-worker image: rayproject/ray:2.56.1-gpu resources: limits: { nvidia.com/gpu: "1" } requests: { cpu: "4", memory: "16Gi", nvidia.com/gpu: "1" }

適用して、サービスの準備が整っていることを確認します。

kubectl apply -f my-service.yaml -n my-namespace kubectl get rayservice my-service -n my-namespace

エンドポイントに到達する

Ray Serve はヘッドポッドのポート8000をリッスンします。クラスター内から、KubeRay が 用に作成するサービスにリクエストを送信しますRayService。クイックテストには、 を使用できますkubectl port-forward。

kubectl port-forward svc/ray-service-head-svc 8000:8000 curl http://localhost:8000/

Serve デプロイ API とリクエスト処理については、Ray ドキュメントの「Ray Serve API」を参照してください。