ParsingConfiguration¶
Structure Class¶
ParsingConfiguration
dataclass
¶
Settings for parsing document contents. If you exclude this field, the
default parser converts the contents of each document into text before
splitting it into chunks. Specify the parsing strategy to use in the
parsingStrategy field and include the relevant configuration, or omit
it to use the Amazon Bedrock default parser. For more information, see
Parsing options for your data
source.
Note
If you specify BEDROCK_DATA_AUTOMATION or BEDROCK_FOUNDATION_MODEL
and it fails to parse a file, the Amazon Bedrock default parser will be
used instead.
Attributes¶
bedrock_data_automation_configuration
class-attribute
instance-attribute
¶
bedrock_data_automation_configuration: BedrockDataAutomationConfiguration | None = None
If you specify BEDROCK_DATA_AUTOMATION as the parsing strategy for
ingesting your data source, use this object to modify configurations for
using the Amazon Bedrock Data Automation parser.
bedrock_foundation_model_configuration
class-attribute
instance-attribute
¶
bedrock_foundation_model_configuration: BedrockFoundationModelConfiguration | None = None
If you specify BEDROCK_FOUNDATION_MODEL as the parsing strategy for
ingesting your data source, use this object to modify configurations for
using a foundation model to parse documents.
parsing_strategy
instance-attribute
¶
parsing_strategy: str
The parsing strategy for the data source.
For managed knowledge bases, the strategy that you can select depends on the embedding model that your knowledge base uses:
-
If your knowledge base uses a native multimodal embedding model, specify
MULTI_MODAL_EMBEDDINGS. With this strategy, files are sent directly to the embedding model instead of being parsed into text. This is the only strategy that is supported for these knowledge bases. -
Otherwise, specify
SMART_PARSING.
For more information, see Customize ingestion for managed knowledge bases.