View a markdown version of this page

Guida introduttiva alla valutazione in batch - Amazon Bedrock AgentCore

Guida introduttiva alla valutazione in batch

Questa procedura dettagliata ti porta da un agente distribuito alla valutazione in batch dei risultati utilizzando un agente dell'assistenza clienti di Acme Store. Potrai creare l'agente, distribuirlo, generare sessioni di esempio, eseguire una valutazione in batch e leggere i risultati.

Prima di iniziare

Assicurati di:

  • La AgentCore CLI installata () agentcore --version

  • AWS credenziali con autorizzazioni per e bedrock-agentcore logs

  • Ricerca delle transazioni abilitata in CloudWatch

  • Python 3.10+ (per esempi di boto3)

Per i dettagli completi, consulta Prerequisiti.

Le seguenti costanti vengono utilizzate negli esempi di boto3. Sostituiscili con i tuoi valori dopo aver distribuito l'agente:

REGION = "us-west-2" AGENT_ARN = "arn:aws:bedrock-agentcore:us-west-2:123456789012:runtime/AcmeSupport-abc123" SERVICE_NAME = "AcmeSupport-abc123.DEFAULT" LOG_GROUP = "/aws/bedrock-agentcore/runtimes/AcmeSupport-abc123-DEFAULT"

Fase 1: Creare e distribuire l'agente di esempio

Crea un AgentCore progetto e sostituisci il codice agente predefinito con l'agente dell'assistenza clienti Acme Store. Questo agente dispone di cinque strumenti per la gestione di ordini, resi, spedizioni, sconti ed escalation.

Creazione del progetto

agentcore create --name AcmeSupport --framework Strands --model-provider Bedrock --memory none cd AcmeSupport

Sostituisci il codice dell'agente

Apri app/AcmeSupport/main.py e sostituisci il suo contenuto con quanto segue:

"""Acme Store customer support agent.""" from strands import Agent, tool from strands.models.bedrock import BedrockModel from bedrock_agentcore.runtime import BedrockAgentCoreApp app = BedrockAgentCoreApp() MODEL_ID = "global.anthropic.claude-sonnet-4-6" SYSTEM_PROMPT = ( "You are a helpful customer support assistant for Acme Store. " "Help customers with their orders, returns, and shipping questions." ) @tool def lookup_order(order_id: str) -> str: """Look up an order by ID and return its status, item, and delivery details.""" orders = { "ORD-1001": { "status": "delivered", "item": "Blue T-Shirt (L)", "delivered": "2026-03-28", "total": "$29.99", }, "ORD-1002": { "status": "in_transit", "item": "Running Shoes (10)", "shipped": "2026-03-30", "est_delivery": "2026-04-05", "total": "$89.99", }, "ORD-1003": { "status": "delayed", "item": "Wireless Headphones", "shipped": "2026-03-25", "est_delivery": "2026-03-29", "days_late": 5, "total": "$59.99", }, "ORD-1004": { "status": "processing", "item": "Yoga Mat", "ordered": "2026-04-02", "total": "$34.99", }, "ORD-1005": { "status": "delivered", "item": "Coffee Maker", "delivered": "2026-03-20", "total": "$149.99", }, } return str(orders.get(order_id, {"error": f"Order {order_id} not found"})) @tool def initiate_return(order_id: str, reason: str) -> str: """Initiate a return for an order. Sends a return label to the customer.""" return ( f"Return initiated for {order_id}. Reason: {reason}. " "Return label sent to customer email. Please ship within 14 days." ) @tool def check_shipping_status(order_id: str) -> str: """Check detailed shipping status including carrier location and delays.""" statuses = { "ORD-1002": ( "Package is with carrier, currently in Portland OR. " "On schedule for April 5." ), "ORD-1003": ( "Package delayed at distribution center in Memphis TN. " "Original delivery was March 29. Now 5 days late. " "Acme Store policy: orders delayed 3+ days qualify for 15% discount." ), } return statuses.get(order_id, f"No active shipment found for {order_id}.") @tool def apply_discount(order_id: str, discount_percent: int, reason: str) -> str: """Apply a percentage discount to an order and issue a refund.""" return ( f"Applied {discount_percent}% discount to {order_id}. " f"Reason: {reason}. Refund will appear in 3-5 business days." ) @tool def escalate_to_human(reason: str) -> str: """Escalate the conversation to a human support agent.""" return ( f"Escalated to human agent. Reason: {reason}. " "Estimated wait time: 3 minutes." ) agent = Agent( model=BedrockModel(model_id=MODEL_ID), tools=[lookup_order, initiate_return, check_shipping_status, apply_discount, escalate_to_human], system_prompt=SYSTEM_PROMPT, ) @app.entrypoint def invoke(payload, context): result = agent(payload.get("prompt", "Hello")) return {"response": str(result)} if __name__ == "__main__": app.run()

Implementa e verifica

agentcore deploy

Dopo la distribuzione, verifica che l'agente sia in esecuzione:

agentcore invoke --prompt "What's the status of order ORD-1001?"

Dovresti vedere una risposta con i dettagli dell'ordine. Annota l'ARN di runtime, il nome del servizio e il gruppo di log diagentcore status --json: ti serviranno per gli esempi di boto3.

Nota

Se hai già un agente distribuito su AgentCore Runtime con l'osservabilità abilitata, salta questo passaggio e usa il tuo agente per il resto della procedura dettagliata.

Fase 2: Generazione di sessioni di esempio

Richiama l'agente con vari prompt per creare sessioni di valutazione. Queste istruzioni coprono diversi scenari: ricerche di ordini, resi, ritardi di spedizione, richieste di discount e interazioni con più strumenti.

Esempio
AgentCore CLI
agentcore invoke --runtime AcmeSupport --prompt "What's the status of my order ORD-1001?" agentcore invoke --runtime AcmeSupport --prompt "I need to return order ORD-1001, the shirt doesn't fit." agentcore invoke --runtime AcmeSupport --prompt "What's the shipping status on ORD-1002?" agentcore invoke --runtime AcmeSupport --prompt "My order ORD-1003 is delayed, can you help?" agentcore invoke --runtime AcmeSupport --prompt "I'd like to check on order ORD-1004 please." agentcore invoke --runtime AcmeSupport --prompt "Can you look up order ORD-1005 for me?" agentcore invoke --runtime AcmeSupport --prompt "I want to return the coffee maker from order ORD-1005, it's defective." agentcore invoke --runtime AcmeSupport --prompt "Where is my order ORD-1002? It should have arrived by now." agentcore invoke --runtime AcmeSupport --prompt "ORD-1003 is really late, I want a discount." agentcore invoke --runtime AcmeSupport --prompt "Can you check order ORD-1001 and tell me when it was delivered?"
AWS SDK (boto3)
import boto3 import json import uuid client = boto3.client("bedrock-agentcore", region_name=REGION) prompts = [ "What's the status of my order ORD-1001?", "I need to return order ORD-1001, the shirt doesn't fit.", "What's the shipping status on ORD-1002?", "My order ORD-1003 is delayed, can you help?", "I'd like to check on order ORD-1004 please.", "Can you look up order ORD-1005 for me?", "I want to return the coffee maker from order ORD-1005, it's defective.", "Where is my order ORD-1002? It should have arrived by now.", "ORD-1003 is really late, I want a discount.", "Can you check order ORD-1001 and tell me when it was delivered?", ] for i, prompt in enumerate(prompts): session_id = f"acme-eval-{uuid.uuid4().hex[:12]}" print(f"[{i+1}/10] {prompt[:60]}...") response = client.invoke_agent_runtime( agentRuntimeArn=AGENT_ARN, runtimeSessionId=session_id, payload=json.dumps({"prompt": prompt}).encode(), ) response_body = response["response"].read() print(f" Done (session: {session_id})") print("\nAll sessions created.")

Attendi 2-3 minuti dopo l'ultima chiamata per acquisire la telemetria prima di procedere. CloudWatch

Fase 3: Eseguire la valutazione in batch

Avvia una valutazione in batch per assegnare un punteggio a tutte le sessioni recenti. Il servizio rileva le sessioni dai CloudWatch registri, esegue ogni valutatore per ogni sessione e restituisce risultati aggregati.

Esempio
AgentCore CLI
agentcore run batch-evaluation \ --runtime AcmeSupport \ --evaluator Builtin.GoalSuccessRate Builtin.Helpfulness Builtin.Faithfulness \ --wait

Per impostazione predefinita, agentcore run batch-evaluation avvia il processo e lo ritorna immediatamente (senza bloccarlo). Passa --wait al blocco finché il lavoro non raggiunge lo stato terminale. Con--wait, la CLI risolve il gruppo di CloudWatch log e il nome del servizio dalla configurazione del progetto, avvia il lavoro, lo blocca fino a raggiungere lo stato terminale e quindi stampa i punteggi medi per valutatore:

Batch evaluation completed: acme-eval-a1b2c3d4

Sessions: 10 completed, 0 failed, 10 total

Evaluator                           Avg Score
─────────────────────────────────────────────
Builtin.GoalSuccessRate             0.7200
Builtin.Helpfulness                 0.8100
Builtin.Faithfulness                0.8500

Results saved to .cli/jobs/batch-eval-results/

Aggiungilo --json per emettere risultati leggibili dalla macchina (incluso il programma di valutazione individualeaverageScore) per la creazione di script batchEvaluationId e per etichettare l'esecuzione in modo da poter confrontare i risultati tra le esecuzioni. -n <name> Ad esempio:

agentcore run batch-evaluation \ --runtime AcmeSupport \ --evaluator Builtin.GoalSuccessRate Builtin.Helpfulness Builtin.Faithfulness \ -n acme_baseline \ --wait
AWS SDK (boto3)
import boto3 import uuid import time import json eval_client = boto3.client("bedrock-agentcore", region_name=REGION) # Start the batch evaluation response = eval_client.start_batch_evaluation( batchEvaluationName=f"acme_baseline_{uuid.uuid4().hex[:8]}", evaluators=[ {"evaluatorId": "Builtin.GoalSuccessRate"}, {"evaluatorId": "Builtin.Helpfulness"}, {"evaluatorId": "Builtin.Faithfulness"}, ], dataSourceConfig={ "cloudWatchLogs": { "serviceNames": [SERVICE_NAME], "logGroupNames": [LOG_GROUP], } }, clientToken=str(uuid.uuid4()), ) batch_eval_id = response["batchEvaluationId"] print(f"Started: {batch_eval_id}") # Poll until complete while True: result = eval_client.get_batch_evaluation(batchEvaluationId=batch_eval_id) status = result["status"] print(f"Status: {status}") if status in ("COMPLETED", "COMPLETED_WITH_ERRORS", "FAILED", "STOPPED"): break time.sleep(30) print(json.dumps(result, indent=4, default=str))

Fase 4: Leggi i dettagli per sessione

I punteggi aggregati forniscono il quadro generale. Per visualizzare i punteggi per turno e per valutatore per le singole sessioni, utilizza i comandi di visualizzazione CLI integrati o leggi gli eventi di valutazione direttamente da Logs. CloudWatch

Esempio
AgentCore CLI

La CLI fornisce comandi di prima classe per visualizzare i processi di valutazione in batch completati e i relativi risultati. Visualizza un lavoro specifico in base all'ID del processo di valutazione in batch o elenca i lavori passati:

# View a batch evaluation job and its results agentcore view batch-evaluation acme-eval-a1b2c3d4 # List batch evaluation jobs agentcore batch-evaluations history

Questi comandi vengono eseguiti in modo interattivo quando non vengono forniti flag. Aggiungi, ad esempio, --json per un output non interattivo e leggibile dalla macchina. agentcore view batch-evaluation acme-eval-a1b2c3d4 --json

AWS SDK (boto3)
# Get the output location from the batch evaluation result output = result["outputConfig"]["cloudWatchConfig"] log_group = output["logGroupName"] log_stream = output["logStreamName"] # Read the events logs_client = boto3.client("logs", region_name=REGION) response = logs_client.get_log_events( logGroupName=log_group, logStreamName=log_stream, ) for event in response["events"]: event_attrs = json.loads(event["message"]).get("attributes", {}) print(f"Score: {event_attrs.get('gen_ai.evaluation.score.value')}") print(f"Label: {event_attrs.get('gen_ai.evaluation.score.label')}") print(f"Explanation: {event_attrs.get('gen_ai.evaluation.explanation', '')[:200]}") print()

Fasi successive