在 Amazon Nova 2 上为 SFT 准备数据
Amazon Nova 2 上的 SFT 支持文本、图像、视频和文档理解以及工具调用,无论是否支持推理。本页介绍为 Amazon Nova 2 理解模型准备 SFT 训练数据的限制、支持的格式和最佳实践。
提示
要在开始训练作业之前验证您的数据集格式,请参阅 验证工具。
数据格式
Amazon Nova 2 SFT 数据采用与 Amazon Nova 1 相同的 Converse API 格式,并新增了可选的推理内容字段。
JSONL 训练文件中的每一行都是一个 JSON 对象,包含以下顶级字段。展开一个部分以了解更多信息:
必需。messages 字段是一个消息对象数组,每个对象定义对话中的一个轮次。消息对象包含以下字段:
-
角色:必填项。定义消息是来自
user(发送给模型的提示)还是assistant(模型响应)。第一个轮次必须是user,最后一个轮次必须是assistant,且轮次必须交替出现。 -
内容:必填项。本轮次的内容块数组。
content 字段映射到一个内容块数组。Amazon Nova 2 SFT 数据支持以下内容块:
一个可选数组,用于定义系统提示 — 关于模型应该执行的任务或应该采用的角色的说明或上下文。在训练和推理期间使用相同的系统提示,以获得最佳结果。
"system": [ { "text": "You are a helpful assistant." } ]
一个可选对象,定义模型在对话期间可以使用的工具。每个工具都使用名称、描述和其输入参数的 JSON Schema 进行定义。
"toolConfig": { "tools": [ { "toolSpec": { "name": "tool-name", "description": "tool-description", "inputSchema": { "json": { "type": "object", "properties": { "param": { "type": "string", "description": "param-description" } }, "required": ["param"] } } } } ] }
必需。标识架构版本的字符串字段。可以是任何字符串值。
"schemaVersion": "bedrock-conversation-2024"
验证数据
在提交训练作业之前,请验证数据集,以尽早发现格式问题。有关可用的验证工具,请参阅 验证工具。
示例输入
以下是完整的 JSON 对象示例,展示了如何针对不同的模式组合各字段和内容块。
支持的功能
下表比较了各 Nova 模型版本对 SFT 的特征支持情况。
通用理解/文本理解
本部分总结了为 Amazon Nova 2 训练数据准备 SFT 的一般限制条件。
约束:
| 约束 | 说明 |
|---|---|
| 数据集格式 | JSONL(每行一个 JSON 对象)。文件名只能包含字母数字字符、下划线、连字符、斜杠和句点。 |
| 最小样本数 | 8 |
| 最大样本数 | 20k |
| 上下文长度 | 32k |
| 数据集同构性 | 数据集不能混合不同的媒体模态。可使用图像配合文本、视频配合文本,或文档配合文本,但不能将多种组合混用。 |
| 保留关键字 | User:、Bot:、Assistant:、System:、<image>、<video>、[EOS]。包含这些关键字的提示将导致训练作业失败。请用含义相似的其他关键字替代它们。 |
最佳实践:
微调所需的最小数据量取决于任务(即是复杂任务还是简单任务),但建议至少为希望模型学习的每项任务提供 200 个数据样本。
建议在训练和推理期间,在零样本设置中使用经过优化的提示,以便获得最佳结果。
质量优先于数量。一般而言,几百个高质量、一致的示例优于数千个杂乱或相互矛盾的示例。
示例输入
schemaVersion可以是任何字符串值支持的角色为
user和assistant。(可选)system轮次可以是客户提供的自定义系统提示。messages中的第一轮应始终以"role": "user"开头。最后一轮是机器人的响应,以"role": "assistant"表示。
{ "schemaVersion": "bedrock-conversation-2024", "system": [ { "text": "You are a digital assistant with a friendly personality" } ], "messages": [ { "role": "user", "content": [ { "text": "What country is right next to Australia?" } ] }, { "role": "assistant", "content": [ { "text": "The closest country is New Zealand" } ] } ] }
图像理解
SFT 支持基于图像任务的训练,使模型能够学习如何分析并回答与图像内容相关的问题。
约束:
| 约束 | 说明 |
|---|---|
| 支持的格式 | PNG、JPEG、GIF、WebP |
| 每个样本的最大图像数 | 10 |
| 图像文件最大大小 | 10 MB |
| 数据集同构性 | 单个样本可包含图像与文本,但不可将图像与其他模态(视频、文档)混用。 |
| S3 位置 | image.source.s3Location.uri 必须与您的数据集位于同一 Amazon S3 存储桶中。例如,若数据集位于 s3://amzn-s3-demo-bucket/train/train.jsonl 中,则图像或视频必须位于 s3://amzn-s3-demo-bucket 中 |
最佳实践:
确保图像质量高且与任务相关。
提供覆盖不同图像类型与问题格式的多样化示例。
包含的问题应明确指向图像内容中的特定内容。
示例输入
{ "schemaVersion": "bedrock-conversation-2024", "system": [ { "text": "You are a helpful assistant." } ], "messages": [ { "role": "user", "content": [ { "image": { "format": "jpeg", "source": { "s3Location": { "uri": "s3://your-bucket/your-path/your-image.jpg", "bucketOwner": "your-aws-account-id" } } } }, { "text": "Which country is highlighted in the image?" } ] }, { "role": "assistant", "content": [ { "text": "The highlighted country is New Zealand" } ] } ] }
视频理解
SFT 支持基于视频任务的训练,使模型能够学习如何分析并回答与视频内容相关的问题。
约束:
| 约束 | 说明 |
|---|---|
| 支持的格式 | MOV、MKV、MP4、WebM |
| 每个样本的最大视频数 | 1 |
| 视频文件最大大小 | 50 MB |
| 视频最长时长 | 15 分钟 |
| 数据集同构性 | 单个样本可包含视频与文本,但不可将视频与其他模态(图像、文档)混用。 |
| S3 位置 | video.source.s3Location.uri 必须与您的数据集位于同一 Amazon S3 存储桶中。例如,若数据集位于 s3://amzn-s3-demo-bucket/train/train.jsonl 中,则视频必须位于 s3://amzn-s3-demo-bucket 中 |
最佳实践:
保持视频简洁明了,聚焦与任务相关的内容。
确保视频画质清晰,便于模型提取有效信息。
提出的问题应明确指向视频中的特定内容。
提供覆盖不同视频类型与问题格式的多样化样本。
示例输入
{ "schemaVersion": "bedrock-conversation-2024", "messages": [ { "role": "user", "content": [ { "text": "What are the ways in which a customer can experience issues during checkout on Amazon?" }, { "video": { "format": "mp4", "source": { "s3Location": { "uri": "s3://my-bucket-name/path/to/videos/customer_service_debugging.mp4", "bucketOwner": "123456789012" } } } } ] }, { "role": "assistant", "content": [ { "text": "Customers can experience issues with 1. Data entry, 2. Payment methods, 3. Connectivity while placing the order. Which one would you like to dive into?" } ] } ] }
文档理解
SFT 支持基于文档任务的训练,使模型能够学习如何分析并回答与 PDF 文档相关的问题。
约束:
| 约束 | 说明 |
|---|---|
| 支持的格式 | |
| 最大文档大小 | 10 MB |
| 数据集同构性 | 单个样本可包含文档与文本,但不可将文档与其他模态(图像、视频)混用。 |
| S3 位置 | document.source.s3Location.uri 必须与您的数据集位于同一 Amazon S3 存储桶中。例如,若数据集位于 s3://amzn-s3-demo-bucket/train/train.jsonl 中,则文档必须位于 s3://amzn-s3-demo-bucket 中 |
最佳实践:
确保文档格式清晰,文本可正常提取。
提供覆盖不同文档类型与问题格式的多样化样本。
包含推理内容,帮助模型学习文档分析模式。
示例输入
{ "schemaVersion": "bedrock-conversation-2024", "messages": [ { "role": "user", "content": [ { "text": "What are the ways in which a customer can experience issues during checkout on Amazon?" }, { "document": { "format": "pdf", "source": { "s3Location": { "uri": "s3://my-bucket-name/path/to/documents/customer_service_debugging.pdf", "bucketOwner": "123456789012" } } } } ] }, { "role": "assistant", "content": [ { "text": "Customers can experience issues with 1. Data entry, 2. Payment methods, 3. Connectivity while placing the order. Which one would you like to dive into?" } ] } ] }
工具调用
SFT 支持基于工具调用模式进行模型训练,使模型能够学习何时以及如何调用外部工具或函数。
约束:
| 约束 | 说明 |
|---|---|
| 支持的格式 | ToolResult 内容的文本或 JSON 格式 |
| ToolUse 放置位置 | ToolUse 只能出现在助手轮次中 |
| ToolResult 放置 | ToolResult 只能出现在用户轮次中 |
| inputSchema 格式 | toolSpec 中的 inputSchema 必须是有效的 JSON 架构对象 |
| toolUseId 匹配 | 每个 ToolResult 必须引用前序助手轮次 ToolUse 中的有效 toolUseId,且每个 toolUseId 在单次对话中仅可使用一次 |
最佳实践:
确保工具定义在所有训练样本中保持一致。
模型将从所提供的示例中学习工具调用模式。
请包含多样化的示例,说明何时应使用每个工具以及何时不应使用工具。
示例输入
{ "schemaVersion": "bedrock-conversation-2024", "system": [ { "text": "You are an expert in composing function calls." } ], "toolConfig": { "tools": [ { "toolSpec": { "name": "getItemAvailability", "description": "Retrieve whether an item is available in a given location", "inputSchema": { "json": { "type": "object", "properties": { "zipcode": { "type": "string", "description": "The zipcode of the location to check in" }, "quantity": { "type": "integer", "description": "The number of items to check availability for" }, "item_id": { "type": "string", "description": "The ASIN of item to check availability for" } }, "required": ["item_id", "zipcode"] } } } } ] }, "messages": [ { "role": "user", "content": [ { "text": "I need to check whether there are twenty pieces of the following item available. Here is the item ASIN on Amazon: id-123. Please check for the zipcode 94086" } ] }, { "role": "assistant", "content": [ { "toolUse": { "toolUseId": "getItemAvailability_0", "name": "getItemAvailability", "input": { "zipcode": "94086", "quantity": 20, "item_id": "id-123" } } } ] }, { "role": "user", "content": [ { "toolResult": { "toolUseId": "getItemAvailability_0", "content": [ { "text": "[{\"name\": \"getItemAvailability\", \"results\": {\"availability\": true}}]" } ] } } ] }, { "role": "assistant", "content": [ { "text": "Yes, there are twenty pieces of item id-123 available at 94086. Would you like to place an order or know the total cost?" } ] } ] }
推理
推理内容(亦称思维链)会记录模型在生成最终答案前的中间思考步骤。
约束:
| 约束 | 说明 |
|---|---|
| 支持的格式 | 仅文本。不支持基于图像的推理内容。 |
| 放置位置 | 仅限助手轮次,通过 reasoningContent 字段指定。 |
| 格式设置 | 使用纯文本。除非任务明确要求,否则避免使用 <thinking> 和 </thinking> 等标记标签。 |
最佳实践:
高质量的推理内容应包括中间思路、逻辑推断、逐步解决问题的方法,以及步骤与结论之间的明确关联。
您可在多轮对话的多个助手轮次中添加
reasoningContent。如果数据集缺失推理轨迹,可借助 Nova Premier 等具备推理能力的模型来生成。
示例输入
{ "schemaVersion": "bedrock-conversation-2024", "system": [ { "text": "You are a digital assistant with a friendly personality" } ], "messages": [ { "role": "user", "content": [ { "text": "What country is right next to Australia?" } ] }, { "role": "assistant", "content": [ { "reasoningContent": { "reasoningText": { "text": "I need to use my world knowledge of geography to answer this question" } } }, { "text": "The closest country to Australia is New Zealand, located to the southeast across the Tasman Sea." } ] } ] }
附加注释:
损失计算方式:
包含推理内容:训练损失同时计入推理词元和最终输出词元。
不含推理内容:训练损失仅基于最终输出词元计算。
当您的训练数据包含推理词元、您希望模型在生成最终输出之前生成思维词元,或需要提高复杂推理任务的性能时,请在训练配置中设置 reasoning_enabled: true。
当您的训练数据没有推理词元、您正在训练无法从显式推理步骤中获益的直接任务,或者您希望优化速度并减少词元的使用时,设置 reasoning_enabled: false。
允许在 reasoning_enabled = true 的情况下,使用非推理数据集训练 Nova。但是,这样做可能会导致模型丧失推理能力,因为 Nova 主要学习数据中呈现的应答方式,而非执行推理过程。一般建议:使用推理数据集时,训练与推理均启用推理;使用非推理数据集时,两者均关闭推理。
设计有效的训练示例
您的训练数据应会表现出您希望模型表现的行为。SFT 仅传授模型如何响应,而不会传授需要知道的知识。如果您发现自己创建的训练示例主要是为了注入事实性知识(例如,“错误代码 E-45 是什么意思?”,答案为“E-45 表示传感器超时”),则应考虑 RAG 或提示工程是否能够在不进行微调的情况下获得相同的结果。
将源数据转换为对话格式时,应遵循以下原则:
-
从真实用户查询出发。编写能够反映最终用户在生产中实际向模型发出提示的方式的用户轮次。避免不能反映真实使用模式的人造问题或过度简化的问题。
-
编写黄金标准的助手响应。每个助手轮次都应是您希望模型产生的理想响应——在所有示例中都正确、格式恰当、语气和结构一致。
-
策略性地使用系统提示。将输出格式、角色和领域上下文等永久说明放入系统轮次。在训练和推理期间使用相同的系统提示,以获得最佳结果。
-
确保每个示例专注于一项技能。如果需要模型执行分类、回答问题和进行总结,应为每项技能创建单独的示例集,而不是将多个技能组合到单个示例中。
-
质量的重要性高于数量。一般而言,几百个高质量、一致的示例优于数千个杂乱或相互矛盾的示例。在训练之前删除重复数据并检查数据。
常用数据转换模式
以下模式演示了如何将常用的企业数据格式转换为有效的 SFT 训练示例。
模式 1:将结构化数据转换为分类或查询
当源数据是由具有诸如 issue、cause 和 resolution 之类字段的记录组成的表或数据库时,则将每条具有相关输入字段的记录作为用户问题,并将具有目标输出字段的记录作为助手的响应。
// train.jsonl { "schemaVersion": "bedrock-conversation-2024", "system": [ { "text": "You are a technical support assistant. Given an issue description, identify the most likely root cause and recommended resolution." } ], "messages": [ { "role": "user", "content": [ { "text": "A user reports: 'Application crashes on startup after the latest update.' What is the likely cause and resolution?" } ] }, { "role": "assistant", "content": [ { "text": "Cause: Incompatible configuration file from the previous version.\nResolution: Delete the local configuration cache and restart the application to regenerate default settings." } ] } ] }
模式 2:将层次结构或分类法数据转换为结构化问答
当源数据是诸如产品类别树或分类系统之类的代码层次结构或分类法数据时,则将生成要求模型浏览层次结构的问题。在系统提示或用户轮次中包含分类上下文。
// train.jsonl { "schemaVersion": "bedrock-conversation-2024", "system": [ { "text": "You are a product catalog assistant. Answer questions about product categories and classifications accurately." } ], "messages": [ { "role": "user", "content": [ { "text": "What top-level category does product code 'EL-4520' belong to, and what subcategories does it include?" } ] }, { "role": "assistant", "content": [ { "text": "Product code EL-4520 belongs to the Electronics category. It includes the subcategories: Audio Equipment, Portable Speakers, and Wireless Accessories." } ] } ] }
模式 3:从流程文档转换为指令性对话
当源数据由标准操作程序、运行手册或分步指南组成时,则会将流程转换为以任务为导向的问题。当用户询问如何执行某个任务时,助手以一致的格式提供操作步骤。
// train.jsonl { "schemaVersion": "bedrock-conversation-2024", "system": [ { "text": "You are an IT operations assistant. Provide step-by-step instructions based on approved runbooks and standard operating procedures." } ], "messages": [ { "role": "user", "content": [ { "text": "How do I reset a user's multi-factor authentication when they have lost access to their device?" } ] }, { "role": "assistant", "content": [ { "text": "Follow these steps to reset a user's multi-factor authentication:\n1. Verify the user's identity using the approved verification procedure.\n2. Disable the existing MFA configuration in the identity management console.\n3. Generate a new temporary access code and send it to the user's verified email.\n4. Instruct the user to log in with the temporary code and re-enroll their new device.\n5. Confirm the new MFA device is working by requesting a test authentication.\n6. Document the reset in the support ticket." } ] } ] }
注意
在模式 1 和 2 中,如果模型的预训练数据中还没有事实内容(例如代码、类别或解决方法),可考虑在推理时使用 RAG 提供这些信息,而不是完全依靠 SFT 来记住这些信息。SFT 对于向模型传授响应格式和推理模式最为有效,而 RAG 则负责事实基础。