View a markdown version of this page

在 AgentCore 网关中使用采样 - 亚马逊基岩 AgentCore

本文属于机器翻译版本。若本译文内容与英语原文存在差异,则一律以英文原文为准。

在 AgentCore 网关中使用采样

采样是一项 MCP 功能,它允许 MCP 服务器在工具调用期间向客户端请求 LLM 完成。这使服务器无需直接访问语言模型即可利用 AI 功能——客户端处理模型调用并返回结果。 AgentCore Gateway 将来自 MCP 服务器目标的采样请求转发到您的客户端,将请求替换为网关生id成的标识符。

先决条件

要在网关中使用采样,请执行以下操作:

  • 启用会话(2025-11-25 及更早版本)— 采样需要会话支持。请参阅在网关上使用 MCP 会话。对于版本2026-07-28及更高版本,您无需sessionConfiguration添加到网关,因为这些版本是无状态的。

  • 启用响应流(2025-11-25 及更早版本)— 在开放连接期间,采样请求以 SSE 区块的形式发送。streamingConfiguration.enableResponseStreaming在网关true中设置为protocolConfiguration.mcp。对于版本2026-07-28及更高版本,您无需启用响应流。这些版本通过多往返请求 (MRTR) 模式提供采样,而不是响应流上服务器发起的请求。有关更多信息,请参阅 “模型上下文协议” 文档中的多次往返请求。

  • MCP 服务器目标类型 -采样请求源自 MCP 服务器目标。

  • 客户端声明采样能力 -客户端必须声明支持采样,网关才能转发采样请求。对于版本2025-11-25及更早版本,客户端会在initialize请求中声明此支持。对于版本2026-07-28及更高版本,客户端在_meta字段 (io.modelcontextprotocol/clientCapabilities) 中为每个请求声明它。

采样的工作原理

当 MCP 服务器目标在工具执行期间需要 LLM 完成时,它会发送请求。sampling/createMessage网关将此请求作为 SSE 事件转发给客户端,取代请求id。客户端调用其语言模型并将结果发送回网关,网关将其转发给目标。

注意

此处描述的流程适用于版本2025-11-25及更早版本,其中服务器sampling/createMessage作为服务器启动的请求在开放的 SSE 流上发送。对于版本2026-07-28及更高版本,采样改为使用多往返请求 (MRTR) 模式。服务器返回一个resultType设置为的临时结果input_required。然后,客户端在重试原始请求时提供完成信息。有关更多信息,请参阅 “模型上下文协议” 文档中的多次往返请求。

采样请求包括:

  • messages— 要发送给模型的对话消息。

  • modelPreferences— 有关所需模型功能(智能、速度、成本)的可选提示。

  • systemPrompt— 模型的可选系统提示。

  • maxTokens— 要生成的最大令牌数。

客户回复如下:

  • model— 使用的模型。

  • role— 永远assistant。

  • content— 生成的内容(文本或图像)。

注意

客户可以完全控制使用哪种模型以及如何处理请求。服务器modelPreferences是提示,不是要求。客户端也可以根据自己的策略修改或拒绝请求。

采样流程

  1. 客户端发送带有Mcp-Session-Id标头的tools/call请求。

  2. 网关将工具调用转发到 MCP 服务器目标。

  3. 目标打开 SSE 直播并发送sampling/createMessage请求。

  4. Gateway 将采样请求作为 SSE 事件转发给客户端,取代请求id。

  5. 客户端使用提供的消息调用其语言模型。

  6. 客户端使用相同Mcp-Session-Id和id来自网关请求的采样结果发送新请求。

  7. 网关将结果转发给 MCP 服务器目标。

  8. 目标继续处理并返回最终的刀具结果。

  9. Gateway 将最终结果转发给客户端并关闭直播。

MCP 服务器目标开发人员指南

重要

发送采样请求的 MCP 服务器目标应将采样调用包装在 try-catch 块中,并处理客户端不支持采样的情况。如果网关的客户端未声明采样能力,则网关不会将其声明给目标。如果目标无论如何都发送采样请求,则网关会向目标返回-32601(未找到方法)错误。

当采样不可用时,服务器应实现备用路径(例如使用内置模型或跳过该 AI-assisted 步骤)。

保护请求状态(版本 2026-07-28 及更高版本)

在版本2026-07-28及更高版本中,采样使用多往返请求 (MRTR) 模式,该模式在您的客户端和 MCP 服务器目标requestState之间呈不透明状态。保护这一价值是一项共同的责任:网关在不存储的情况下对其进行授权和转发,而您的 MCP 服务器目标必须对其进行验证并防止一个用户重播另一个用户的请求状态。有关完全分担责任模型和您的 MCP 服务器必须遵循的保护指南,请参阅 MCP 服务器目标注意事项中的保护请求状态以进行引用和采样。

错误处理

场景 错误 说明

当没有待处理的采样请求时,客户端会发送采样响应

JSON-RPC -32600(请求无效)

未找到该会话的匹配采样请求。

客户端发送的采样响应与待处理的请求不匹配 id

JSON-RPC -32600(请求无效)

id必须与sampling/createMessage请求中网关发送的相匹配。

MCP 服务器发送采样请求,但网关未声明支持

JSON-RPC -32601(未找到方法)

返回到 MCP 服务器目标。参阅故障排除。

问题排查

错误:“调用工具'sample_tool'时出错:未找到方法:” sampling/createMessage

当 MCP 服务器目标发送采样请求但网关的客户端未声明采样能力时,就会出现此错误。对于版本2025-11-25及更早版本,客户端将在期间声明此功能initialize。对于版本2026-07-28及更高版本,客户端会在_meta字段中为每个请求声明该版本。网关向目标返回-32601(未找到方法)错误。目标可能会将其作为工具执行错误返回给客户端。

要解决这个问题,请执行以下操作:

  • 如果您是 MCP 服务器开发人员:在采样调用周围添加错误处理。在不支持采样时实现备用路径:

    重要

    你必须在create_message通话related_request_id=ctx.request_context.request_id中包括在内。这是网关正确地将采样请求与原始工具调用关联所必需的。没有它,采样将不起作用。

    try: result = await ctx.session.create_message( messages=[{"role": "user", "content": {"type": "text", "text": "Summarize this document"}}], max_tokens=500, related_request_id=ctx.request_context.request_id, ) except Exception as e: # Fallback when client doesn't support sampling logger.warning(f"Sampling not supported: {e}") result = fallback_summarization(document)
  • 如果您是网关客户端开发者:对于版本2025-11-25及更早版本,请确保您的客户端在此期间声明采样能力initialize。对于版本2026-07-28及更高版本,请在_meta字段 (io.modelcontextprotocol/clientCapabilities) 中为每个请求声明该版本。以下示例显示了initialize声明:

    { "capabilities": { "sampling": {} } }

代码示例

注意

LangGraph MCP 客户端 (langchain-mcp-adapters) 和 Strands MCP 客户端目前不支持采样。使用如下所示的 MCP 客户端方法来处理来自网关的采样请求。

例
Python requests package (2025-11-25 and earlier)

在这些版本中,客户端在期间声明采样能力initialize,采样请求作为sampling/createMessage请求到达开放的 SSE 流。将标MCP-Protocol-Version头设置为您的网关支持的版本。

import requests import json import sseclient gateway_url = "https://mygateway-abcdefghij.gateway.bedrock-agentcore.us-west-2.amazonaws.com/mcp" headers = { "Content-Type": "application/json", "Accept": "text/event-stream", "Authorization": "Bearer YOUR_ACCESS_TOKEN" } # Step 1: Initialize with sampling capability init_response = requests.post(gateway_url, headers=headers, json={ "jsonrpc": "2.0", "id": "init-request", "method": "initialize", "params": { "protocolVersion": "2025-06-18", "capabilities": {"sampling": {}}, "clientInfo": {"name": "my-agent", "version": "1.0.0"} } }) session_id = init_response.headers["Mcp-Session-Id"] headers["Mcp-Session-Id"] = session_id headers["MCP-Protocol-Version"] = "2025-06-18" # Step 2: Call tool (streaming response) response = requests.post(gateway_url, headers=headers, json={ "jsonrpc": "2.0", "id": "tool-call-1", "method": "tools/call", "params": { "name": "summarizeDocument", "arguments": {"documentId": "doc-789"} } }, stream=True) # Step 3: Process SSE events client = sseclient.SSEClient(response) for event in client.events(): data = json.loads(event.data) if data.get("method") == "sampling/createMessage": sampling_id = data["id"] print(f"Sampling request: {data['params']['messages']}") # Step 4: Invoke your LLM and send result llm_result = invoke_your_model(data["params"]) # Your LLM invocation requests.post(gateway_url, headers=headers, json={ "jsonrpc": "2.0", "id": sampling_id, "result": { "model": "claude-sonnet-4-20250514", "role": "assistant", "content": {"type": "text", "text": llm_result} } }) elif "result" in data: print(f"Tool result: {data['result']}") break
Python requests package (2026-07-28)

在版本上2026-07-28,采样使用多往返请求模式,而不是服务器在 SSE 流上发起的请求。客户端在每个请求中_meta声明采样能力。如果工具需要完成,则响应是包含sampling/createMessage请求inputRequests和不透明input_requiredrequestState的结果。客户端调用其模型并使用新的id、和未修改的重试原始请求。inputResponses requestState不使用会话和initialize握手。您的网关supportedVersions必须包括2026-07-28。

import requests gateway_url = "https://mygateway-abcdefghij.gateway.bedrock-agentcore.us-west-2.amazonaws.com/mcp" META = { "io.modelcontextprotocol/protocolVersion": "2026-07-28", "io.modelcontextprotocol/clientInfo": {"name": "my-agent", "version": "1.0.0"}, "io.modelcontextprotocol/clientCapabilities": {"sampling": {}} } headers = { "Content-Type": "application/json", "Accept": "application/json, text/event-stream", "Authorization": "Bearer YOUR_ACCESS_TOKEN", "MCP-Protocol-Version": "2026-07-28", "Mcp-Method": "tools/call", "Mcp-Name": "summarizeDocument" } arguments = {"documentId": "doc-789"} # Step 1: Call the tool, declaring the sampling capability in _meta response = requests.post(gateway_url, headers=headers, json={ "jsonrpc": "2.0", "id": "tool-call-1", "method": "tools/call", "params": {"name": "summarizeDocument", "arguments": arguments, "_meta": META} }).json() result = response["result"] if result.get("resultType") == "input_required": # Step 2: Fulfill each sampling request by invoking your model input_responses = {} for key, input_request in result.get("inputRequests", {}).items(): params = input_request["params"] print(f"Sampling request: {params['messages']}") llm_result = invoke_your_model(params) # Your LLM invocation input_responses[key] = { "model": "claude-sonnet-4-20250514", "role": "assistant", "content": {"type": "text", "text": llm_result} } # Step 3: Retry the tool call with a new id, the input responses, # and the requestState echoed back unmodified retry_params = {"name": "summarizeDocument", "arguments": arguments, "_meta": META, "inputResponses": input_responses} if "requestState" in result: retry_params["requestState"] = result["requestState"] response = requests.post(gateway_url, headers=headers, json={ "jsonrpc": "2.0", "id": "tool-call-2", "method": "tools/call", "params": retry_params }).json() result = response["result"] print(f"Tool result: {result}")
MCP Client
from mcp import ClientSession from mcp.client.streamable_http import streamablehttp_client import asyncio async def sampling_handler(request): """Handle sampling requests from the server by invoking an LLM.""" messages = request.params.messages llm_response = await invoke_your_model(messages, max_tokens=request.params.maxTokens) return { "model": "claude-sonnet-4-20250514", "role": "assistant", "content": {"type": "text", "text": llm_response} } async def use_sampling(url, token): headers = {"Authorization": f"Bearer {token}"} async with streamablehttp_client(url=url, headers=headers) as ( read_stream, write_stream, _ ): async with ClientSession( read_stream, write_stream, sampling_handler=sampling_handler ) as session: await session.initialize() result = await session.call_tool( name="summarizeDocument", arguments={"documentId": "doc-789"} ) print(f"Tool result: {result}") return result asyncio.run(use_sampling( url="https://mygateway-abcdefghij.gateway.bedrock-agentcore.us-west-2.amazonaws.com/mcp", token="YOUR_ACCESS_TOKEN" ))