

本文為英文版的機器翻譯版本，如內容有任何歧義或不一致之處，概以英文版為準。

# 基準生成式 AI 推論端點
<a name="generative-ai-inference-recommendations-benchmark"></a>

SageMaker AI 基準測試服務會測量 SageMaker AI 端點上託管的大型語言模型 (LLMs) 的效能。它使用 NVIDIA AIPerf 執行基準測試，產生請求延遲、輸送量、首次權杖前所經時間，以及權杖間延遲等指標。

## 先決條件
<a name="generative-ai-inference-recommendations-benchmark-prereqs"></a>

建立基準任務之前，您需要下列項目：
+ 狀態為`InService`託管支援 OpenAI 相容聊天完成 API 之 LLM 的 SageMaker AI 端點
+ 基準輸出的 Amazon S3 儲存貯體
+ IAM 執行角色，授予 SageMaker AI 存取端點和輸出儲存貯體的權限

## 步驟 1：建立基準工作
<a name="generative-ai-inference-recommendations-benchmark-create"></a>

基準任務以特定 SageMaker AI 端點為目標，並參考工作負載組態。

**Python (boto3)**

```
response = client.create_ai_benchmark_job(
    AIBenchmarkJobName="my-benchmark-job",
    BenchmarkTarget={
        "Endpoint": {
            "Identifier": "my-sagemaker-endpoint"
        }
    },
    OutputConfig={
        "S3OutputLocation": "s3://DOC-EXAMPLE-BUCKET/benchmark-results/"
    },
    AIWorkloadConfigIdentifier="my-benchmark-config",
    RoleArn="arn:aws:iam::111122223333:role/ExampleRole",
)
print(response["AIBenchmarkJobArn"])
```

**AWS CLI**

```
aws sagemaker create-ai-benchmark-job \
  --ai-benchmark-job-name "my-benchmark-job" \
  --benchmark-target '{"Endpoint": {"Identifier": "my-sagemaker-endpoint"}}' \
  --output-config '{"S3OutputLocation": "s3://DOC-EXAMPLE-BUCKET/benchmark-results/"}' \
  --ai-workload-config-identifier "my-benchmark-config" \
  --role-arn "arn:aws:iam::111122223333:role/ExampleRole" \
  --region us-west-2
```

如果您的端點透過推論元件託管多個模型，您可以在 的 `InferenceComponents` 參數中指定它們`BenchmarkTarget`。

如果您的端點位於 VPC 中，請將 `NetworkConfig` 參數與您的`VpcConfig`設定一起傳遞，包括安全群組 IDs和子網路。

若要在 SageMaker AI 上使用全受管 MLflow 追蹤基準結果，請將 `MlflowConfig` 物件新增至 `OutputConfig`。如需詳細資訊，請參閱[使用 MLflow 追蹤推論建議和基準結果](generative-ai-inference-recommendations-mlflow.md)。

## 基準推論元件
<a name="generative-ai-inference-recommendations-benchmark-inference-components"></a>

如果您的端點使用*推論元件*而非直接部署模型，您必須在 中指定要進行基準測試的推論元件`BenchmarkTarget`。指定推論元件時，基準測試服務會將請求路由到這些特定元件，而不是端點的預設模型。

在`InferenceComponents`清單中傳遞一或多個推論元件名稱或 ARNs：

**Python (boto3)**

```
response = client.create_ai_benchmark_job(
    AIBenchmarkJobName="my-ic-benchmark",
    BenchmarkTarget={
        "Endpoint": {
            "Identifier": "my-multi-model-endpoint",
            "InferenceComponents": [
                {"Identifier": "my-inference-component-llama"}
            ]
        }
    },
    OutputConfig={
        "S3OutputLocation": "s3://DOC-EXAMPLE-BUCKET/benchmark-results/"
    },
    AIWorkloadConfigIdentifier="my-benchmark-config",
    RoleArn="arn:aws:iam::111122223333:role/ExampleRole",
)
```

**AWS CLI**

```
aws sagemaker create-ai-benchmark-job \
  --ai-benchmark-job-name "my-ic-benchmark" \
  --benchmark-target '{
    "Endpoint": {
      "Identifier": "my-multi-model-endpoint",
      "InferenceComponents": [
        {"Identifier": "my-inference-component-llama"}
      ]
    }
  }' \
  --output-config '{"S3OutputLocation": "s3://DOC-EXAMPLE-BUCKET/benchmark-results/"}' \
  --ai-workload-config-identifier "my-benchmark-config" \
  --role-arn "arn:aws:iam::111122223333:role/ExampleRole" \
  --region us-west-2
```

**注意**  
如果您的端點已針對推論元件設定，但未在基準目標`InferenceComponents`中指定 ，則任務會失敗，並顯示沒有模型直接部署在端點上。對inference-component-based端點進行基準測試時，請務必包含 `InferenceComponents` 參數。

## 多重 LoRA 端點基準測試
<a name="generative-ai-inference-recommendations-benchmark-multi-lora"></a>

若要對提供多個 LoRA 轉接器的端點進行基準測試，請在 中將每個轉接器指定為推論元件`BenchmarkTarget`。您可以選擇性地使用`model_selection_strategy`工作負載參數來控制基準如何在轉接器之間分配請求。如果您未指定策略，則預設為 `round_robin`。

首先，建立工作負載組態。下列範例包含選用`model_selection_strategy`參數：

```
# Create a workload config for multi-LoRA benchmarking
workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "prompt_input_tokens_mean": 550,
        "output_tokens_mean": 150,
        "concurrency": 10,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B",
        "model_selection_strategy": "round_robin"
    },
    "secrets": {
        "hf_token": "arn:aws:secretsmanager:us-west-2:111122223333:secret:my-hf-token-AbCdEf"
    },
    "tooling": {"api_standard": "openai"}
}

import json
client.create_ai_workload_config(
    AIWorkloadConfigName="multi-lora-config",
    WorkloadSpec={"Inline": json.dumps(workload_spec)}
)
```

然後，建立以所有 LoRA 轉接器推論元件為目標的基準工作：

```
response = client.create_ai_benchmark_job(
    AIBenchmarkJobName="multi-lora-benchmark",
    BenchmarkTarget={
        "Endpoint": {
            "Identifier": "my-lora-endpoint",
            "InferenceComponents": [
                {"Identifier": "lora-adapter-customer-support"},
                {"Identifier": "lora-adapter-code-generation"},
                {"Identifier": "lora-adapter-summarization"}
            ]
        }
    },
    OutputConfig={
        "S3OutputLocation": "s3://DOC-EXAMPLE-BUCKET/multi-lora-results/"
    },
    AIWorkloadConfigIdentifier="multi-lora-config",
    RoleArn="arn:aws:iam::111122223333:role/ExampleRole",
)
```

`model_selection_strategy` 參數是選用的，並決定基準工具如何將請求分配到指定的推論元件。有效的值如下：
+ `round_robin` （預設） — 每個轉接器會依序接收請求。第 n 個請求會傳送到 (n 個模組number-of-models)第 個轉接器。
+ `random` — 每個請求都會隨機一致地指派給轉接器。

如果您未指定 `model_selection_strategy`，`round_robin`根據預設基準會使用 。

## 使用合成映像為多模式端點建立基準
<a name="generative-ai-inference-recommendations-benchmark-image"></a>

您可以在工作負載組態中產生合成映像，藉此對視覺語言模型進行基準測試。基準測試服務使用 AIPerf 建立具有可設定維度和格式的影像，然後將它們作為 base64 編碼的承載傳送至您的端點。

下列範例會建立工作負載組態，以使用合成映像對視覺語言模型進行基準測試：

```
import json

workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "image_width_mean": 640,
        "image_height_mean": 480,
        "prompt_input_tokens_mean": 100,
        "output_tokens_mean": 150,
        "concurrency": 8,
        "request_count": 100,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B"
    },
    "secrets": {
        "hf_token": "arn:aws:secretsmanager:us-west-2:111122223333:secret:my-hf-token-AbCdEf"
    }
}

client.create_ai_workload_config(
    AIWorkloadConfigName="image-benchmark-config",
    WorkloadSpec={"Inline": json.dumps(workload_spec)}
)
```

下列參數控制合成映像產生：


| 參數 | Type | 預設 | 說明 | 
| --- | --- | --- | --- | 
| image\_width\_mean | float | 無 | 平均影像寬度，以像素為單位。 | 
| image\_width\_stddev | float | 無 | 影像寬度的標準差。設定為跨請求變更影像維度。 | 
| image\_height\_mean | float | 無 | 平均影像高度，以像素為單位。 | 
| image\_height\_stddev | float | 無 | 影像高度的標準偏差。 | 
| image\_batch\_size | int | 1 | 每個請求的影像數量。 | 
| image\_format | string | png | 映像格式。有效值： png （無損）、 jpeg （遺失、較小的檔案）、 random（每個影像隨機選取 PNG 或 JPEG)。 | 

**可變大小影像**

使用標準差參數產生不同維度的影像，模擬影像大小不同的真實工作負載：

```
workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "image_width_mean": 800,
        "image_width_stddev": 200,
        "image_height_mean": 600,
        "image_height_stddev": 150,
        "image_batch_size": 2,
        "concurrency": 4,
        "request_count": 50,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B"
    }
}
```

## 使用合成影片為多模式端點建立基準
<a name="generative-ai-inference-recommendations-benchmark-video"></a>

您可以在工作負載組態中產生合成影片，藉此對處理影片輸入的多模式模型進行基準測試。基準測試服務使用 AIPerf 的合成影片產生來建立具有可設定解析度、影格速率、持續時間和編碼的影片，然後將它們做為 base64 編碼承載傳送至您的端點。

**注意**  
影片產生預設為停用。您必須在工作負載組態`video_height`中同時指定 `video_width`和 ，才能啟用它。

下列範例會建立工作負載組態，以使用 640×480 解析度的合成影片對多模態模型進行基準測試：

```
import json

workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "video_width": 640,
        "video_height": 480,
        "video_fps": 4,
        "video_duration": 5.0,
        "output_tokens_mean": 150,
        "concurrency": 4,
        "request_count": 50,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B"
    },
    "secrets": {
        "hf_token": "arn:aws:secretsmanager:us-west-2:111122223333:secret:my-hf-token-AbCdEf"
    }
}

client.create_ai_workload_config(
    AIWorkloadConfigName="video-benchmark-config",
    WorkloadSpec={"Inline": json.dumps(workload_spec)}
)
```

### 影片參數
<a name="generative-ai-inference-recommendations-benchmark-video-params"></a>

下列參數控制合成影片產生：


| 參數 | Type | 預設 | 說明 | 
| --- | --- | --- | --- | 
| video\_width | int | 無 | 以像素為單位的影格寬度。必須使用 設定video\_height以啟用影片產生。 | 
| video\_height | int | 無 | 以像素為單位的影格高度。必須使用 設定video\_width以啟用影片產生。 | 
| video\_fps | int | 4 | 每秒影格數。 | 
| video\_duration | float | 5.0 | 以秒為單位的剪輯持續時間。 | 
| video\_batch\_size | int | 1 | 每個請求的影片數量。 | 
| video\_synth\_type | string | moving\_shapes | 合成模式。有效值： moving\_shapes （動畫幾何形狀）、 grid\_clock（具有時鐘動畫的網格）、 noise（隨機像素雜訊）。 | 
| video\_format | string | Webm | 容器格式。有效值：webm。 | 
| video\_codec | string | libvpx-vp9 | 視訊轉碼器。支援的值： libvpx-vp9 (VP9、WebM)。 | 

**注意**  
基準測試服務僅支援使用 WebM 格式的 VP9 編碼。

### 內嵌音軌
<a name="generative-ai-inference-recommendations-benchmark-video-audio"></a>

對於同時處理視訊和音訊的模型，您可以在產生的視訊中嵌入合成音訊軌。音訊預設為停用。`video_audio_num_channels` 設定為 `1`（單聲道） 或 `2`（立體聲） 以啟用它。


| 參數 | Type | 預設 | 說明 | 
| --- | --- | --- | --- | 
| video\_audio\_num\_channels | int | 0 | 0 = 停用，1 = 單聲道，2 = 立體聲。 | 
| video\_audio\_sample\_rate | int | 44100 | 以 Hz (8000–96000) 為單位的取樣率。 | 
| video\_audio\_codec | string | auto | 音訊轉碼器。自動選取 libvorbis WebM 和 MP4 aac的 。您可以使用 aac、 libvorbis或 覆寫 libopus。 | 
| video\_audio\_depth | int | 16 | 每個樣本的位元深度 (8、16、24 或 32)。 | 

### 影片基準測試範例
<a name="generative-ai-inference-recommendations-benchmark-video-examples"></a>

**低解析度影片理解**

```
workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "video_width": 320,
        "video_height": 240,
        "video_fps": 2,
        "video_duration": 3.0,
        "video_synth_type": "moving_shapes",
        "concurrency": 4,
        "request_count": 50,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B"
    }
}
```

**HD 影片基準測試**

```
workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "video_width": 1920,
        "video_height": 1080,
        "video_fps": 8,
        "video_duration": 10.0,
        "concurrency": 2,
        "request_count": 20,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B"
    }
}
```

**多模態模型的音訊視訊**

```
workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "video_width": 640,
        "video_height": 480,
        "video_fps": 4,
        "video_duration": 5.0,
        "video_audio_num_channels": 1,
        "video_audio_sample_rate": 16000,
        "concurrency": 4,
        "request_count": 50,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B"
    }
}
```

**混合文字和影片**

將影片與影片問答或字幕工作負載的文字提示結合在一起：

```
workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "video_width": 640,
        "video_height": 480,
        "video_fps": 4,
        "video_duration": 5.0,
        "prompt_input_tokens_mean": 100,
        "output_tokens_mean": 50,
        "concurrency": 8,
        "request_count": 100,
        "streaming": True,
        "tokenizer": "meta-llama/Llama-3.2-1B"
    }
}
```

### 效能考量
<a name="generative-ai-inference-recommendations-benchmark-video-considerations"></a>
+ 較高的解析度和影格率會增加影片編碼時間和承載大小。對於高輸送量測試，請使用較低的解析度 (320×240 或 640×480)。
+ 具有 WebM 格式的 VP9 (`libvpx-vp9`) 是唯一支援的轉碼器，並為基準化承載提供良好的壓縮。
+ 與視訊串流相比，音訊增加的最低額外負荷。針對以語音為中心的工作負載，使用 16 kHz 的單聲道 (`1`)。

## 步驟 2：監控任務狀態
<a name="generative-ai-inference-recommendations-benchmark-monitor"></a>

輪詢任務狀態，直到達到結束狀態為止。

**Python (boto3)**

```
import time

while True:
    response = client.describe_ai_benchmark_job(
        AIBenchmarkJobName="my-benchmark-job"
    )
    status = response["AIBenchmarkJobStatus"]
    print(f"Status: {status}")
    if status in ("Completed", "Failed", "Stopped"):
        break
    time.sleep(30)

if status == "Completed":
    print(f"Results at: {response['OutputConfig']['S3OutputLocation']}")
elif status == "Failed":
    print(f"Job failed: {response.get('FailureReason', 'unknown')}")
```

**AWS CLI**

```
aws sagemaker describe-ai-benchmark-job \
  --ai-benchmark-job-name "my-benchmark-job" \
  --region us-west-2
```

## 步驟 3：檢閱基準結果
<a name="generative-ai-inference-recommendations-benchmark-results"></a>

基準結果會寫入您指定的 Amazon S3 輸出位置。結果包含下列關鍵指標：

`request_throughput`  
每秒請求數。

`request_latency`  
包含百分位數明細 (P50、P90、P99) End-to-end請求延遲。

`time_to_first_token`  
從提交請求到收到第一個字符的時間。

`inter_token_latency`  
連續輸出字符之間的時間。

`output_token_throughput`  
每秒產生的輸出字符。

每個指標都包含統計摘要：平均值、最小值、最大值、P50、P90、P99 和標準差。

## 基準自訂格式端點
<a name="generative-ai-inference-recommendations-benchmark-custom-format"></a>

如果您的端點使用自訂請求或回應格式 （例如，DJL 自訂處理常式、TensorRT-LLM 原生格式，或其他不支援 OpenAI 聊天完成 API 的服務架構），您可以使用 Jinja2 範本來建立基準，以定義端點的承載形狀。

基準測試服務會使用合成或自訂提示，依請求轉譯範本，然後將承載轉送至 SageMaker AI 端點。您也可以指定 JMESPath 查詢，從端點的回應中擷取產生的文字。

### 建立承載範本
<a name="generative-ai-inference-recommendations-benchmark-custom-format-template"></a>

將端點的請求格式定義為 Jinja2 範本檔案。範本中提供下列變數：


| 變數 | 說明 | 
| --- | --- | 
| text | 第一個文字內容 （合成或從資料集）。 | 
| texts | 所有文字內容的清單。 | 
| model | 模型名稱。 | 
| max\_tokens | 輸出字符限制。 | 
| stream | 是否啟用串流。 | 

使用`|tojson`篩選條件來正確逸出字串值的 JSON。下列範例顯示工具呼叫格式的 DJL 端點範本：

```
{
  "messages": [{"role": "user", "content": {{ text|tojson }}}],
  "tools": [
    {"function": {"name": "Chit_Chat", "description": "casual conversation",
      "parameters": {"type": "object", "properties": {}}}}
  ],
  "max_tokens": {{ max_tokens }}
}
```

將範本檔案上傳至 Amazon S3 （例如 `s3://DOC-EXAMPLE-BUCKET/templates/my_endpoint_template.jinja`)。

### 建立工作負載組態
<a name="generative-ai-inference-recommendations-benchmark-custom-format-config"></a>

建立參考範本檔案的工作負載組態。使用 `extra_inputs` 參數來指定範本路徑和回應擷取查詢。透過`DatasetConfig`頻道將範本檔案交付至基準容器。

```
import json

TEMPLATE_LOCAL_PATH = "/opt/ml/input/data/template/my_endpoint_template.jinja"
RESPONSE_FIELD = "generation_details.generations[0].content"

workload_spec = {
    "benchmark": {"type": "aiperf"},
    "parameters": {
        "extra_inputs": f"payload_template:{TEMPLATE_LOCAL_PATH} response_field:{RESPONSE_FIELD}",
        "tokenizer": "meta-llama/Llama-3.2-1B",
        # ... other benchmark parameters (concurrency, request_count, etc.)
    },
}

response = client.create_ai_workload_config(
    AIWorkloadConfigName="custom-format-config",
    AIWorkloadConfigs={"WorkloadSpec": {"Inline": json.dumps(workload_spec)}},
    DatasetConfig={
        "InputDataConfig": [
            {
                "ChannelName": "template",
                "DataSource": {
                    "S3DataSource": {
                        "S3Uri": "s3://DOC-EXAMPLE-BUCKET/templates/my_endpoint_template.jinja",
                    }
                },
            }
        ]
    },
)
```

`ChannelName` 決定檔案出現在基準容器內的位置。名為 的頻道`template`可在 使用 檔案`/opt/ml/input/data/template/<filename>`。

此`response_field`值是 [JMESPath](https://jmespath.org/) 查詢，可從端點的回應擷取產生的文字。常見的模式包括：
+ `choices[0].message.content` — OpenAI 格式
+ `generation_details.generations[0].content` — DJL 格式
+ `output.text` — 簡單文字回應

如果您省略 `response_field`，基準測試工具會自動偵測回應格式。

### 執行基準測試
<a name="generative-ai-inference-recommendations-benchmark-custom-format-run"></a>

建立以端點為目標的基準任務。服務會自動從 中的 `payload_template`金鑰偵測範本模式，`extra_inputs`並透過適當的代理路徑路由請求。

```
response = client.create_ai_benchmark_job(
    AIBenchmarkJobName="custom-format-benchmark",
    BenchmarkTarget={"Endpoint": {"Identifier": "my-custom-endpoint"}},
    OutputConfig={"S3OutputLocation": "s3://DOC-EXAMPLE-BUCKET/results/"},
    AIWorkloadConfigIdentifier="custom-format-config",
    RoleArn="arn:aws:iam::111122223333:role/ExampleRole",
)
```

### 考量事項
<a name="generative-ai-inference-recommendations-benchmark-custom-format-notes"></a>
+ 範本必須透過 Amazon S3 以檔案形式交付。不支援`extra_inputs`字串中的內嵌 JSON 範本，因為 JSON 中的逗號與參數剖析器衝突。
+ 支援具有單一推論元件的端點。範本模式中不支援具有多個推論元件的端點，因為服務無法判斷從任意承載格式將每個請求路由到哪個元件。
+ 支援串流和非串流端點。在工作負載參數`"streaming": false`中設定 `"streaming": true`或 。

## 關聯基準提示和回應
<a name="generative-ai-inference-recommendations-benchmark-correlate-io"></a>

基準工作完成後，輸出成品會包含每個請求的輸入和輸出資料，您可以將這些資料合併以進行後製處理。這可啟用使用案例，例如品質評估、安全稽核、跨組態的回應比較，以及偵錯非預期的模型行為。

### 基準輸出成品
<a name="generative-ai-inference-recommendations-benchmark-correlate-io-artifacts"></a>

Amazon S3 輸出位置中的`output.tar.gz`封存包含下列與提示-回應相互關聯相關的檔案：

`inputs.json`  
產生對話的完整合成資料集。每個記錄都有 `session_id`和完整的請求承載 （訊息、max\_tokens、模型）。

`outputs.json`  
傳送至端點之請求的依請求回應中繼資料。每個記錄都包含模型的回應文字、每個請求延遲、輸出字符計數，以及`conversation_id`對應回輸入的 。

### 將輸入聯結至輸出
<a name="generative-ai-inference-recommendations-benchmark-correlate-io-join"></a>

使用 中的 `conversation_id` 欄位`outputs.json`和 中的 `session_id` 欄位，將提示與其回應建立關聯`inputs.json`：

```
import json

# Load the artifacts extracted from output.tar.gz in your S3 output location
with open("inputs.json") as f:
    inputs = json.load(f)
with open("outputs.json") as f:
    outputs = json.load(f)

# Build lookup tables and join
inputs_by_id = {rec["session_id"]: rec for rec in inputs["data"]}
outputs_by_id = {rec["conversation_id"]: rec for rec in outputs["data"]}

matched_ids = set(inputs_by_id.keys()) & set(outputs_by_id.keys())
print(f"Matched: {len(matched_ids)} prompt-response pairs")

# Display a correlated sample
for sid in sorted(matched_ids)[:3]:
    in_rec = inputs_by_id[sid]
    out_rec = outputs_by_id[sid]
    prompt = in_rec["payloads"][0]["messages"][0]["content"][:100]
    response = out_rec.get("response_text", "")[:100]
    latency = out_rec["metrics"]["request_latency"]
    print(f"\n[{sid}] Latency: {latency:.0f}ms")
    print(f"  Prompt:   {prompt}...")
    print(f"  Response: {response}...")
```

### 重要說明
<a name="generative-ai-inference-recommendations-benchmark-correlate-io-notes"></a>
+ **聯結金鑰。**在 `conversation_id` 中使用 `outputs.json` 來比對 `session_id`中的 `inputs.json`。不要使用 `session_num`做為位置索引 - 它代表執行順序，這與並行大於 1 時的輸入建立順序不同。
+ **並非所有輸入都有輸出。**`inputs.json` 檔案包含完整產生的資料集集區。當 `request_count` 小於集區大小時，只會將一部分的對話傳送至端點。不相符的輸入是產生但未使用的對話。
+ **輸出結構描述。**中的每個記錄`outputs.json`都包含 `conversation_id`、`response_text`、 `metrics`（使用 `request_latency`和 `output_sequence_length`)，以及計時欄位 (`request_start_ns`、)`request_end_ns`。

## 管理基準測試資源
<a name="generative-ai-inference-recommendations-benchmark-manage"></a>

使用下列操作來管理您的基準工作和工作負載組態。

```
# List benchmark jobs
response = client.list_ai_benchmark_jobs(MaxResults=10)
for job in response["AIBenchmarkJobs"]:
    print(f"{job['AIBenchmarkJobName']} - {job['AIBenchmarkJobStatus']}")

# Stop a running job
client.stop_ai_benchmark_job(
    AIBenchmarkJobName="my-benchmark-job"
)

# Delete a job (must be in a terminal state)
client.delete_ai_benchmark_job(
    AIBenchmarkJobName="my-benchmark-job"
)

# List workload configurations
response = client.list_ai_workload_configs(MaxResults=10)
for config in response["AIWorkloadConfigs"]:
    print(f"{config['AIWorkloadConfigName']} - {config['AIWorkloadConfigArn']}")

# Delete a workload configuration
client.delete_ai_workload_config(
    AIWorkloadConfigName="my-benchmark-config"
)
```