This is the new CloudFormation Template Reference Guide. Please update your bookmarks and links. For help getting started with CloudFormation, see the AWS CloudFormation User Guide.
AWS::Comprehend::EntityRecognizer AugmentedManifestsListItem
An augmented manifest file that provides training data for your custom model. An augmented manifest file is a labeled dataset that is produced by Amazon SageMaker Ground Truth.
Syntax
To declare this entity in your CloudFormation template, use the following syntax:
JSON
{ "AnnotationDataS3Uri" :String, "AttributeNames" :[ String, ... ], "DocumentType" :String, "S3Uri" :String, "SourceDocumentsS3Uri" :String, "Split" :String}
YAML
AnnotationDataS3Uri:StringAttributeNames:- StringDocumentType:StringS3Uri:StringSourceDocumentsS3Uri:StringSplit:String
Properties
AnnotationDataS3Uri-
The S3 prefix to the annotation files that are referred in the augmented manifest file.
Required: No
Type: String
Pattern:
^s3://[a-z0-9][\.\-a-z0-9]{1,61}[a-z0-9](/.*)?$Maximum:
1024Update requires: Replacement
AttributeNames-
The JSON attribute that contains the annotations for your training documents. The number of attribute names that you specify depends on whether your augmented manifest file is the output of a single labeling job or a chained labeling job.
If your file is the output of a single labeling job, specify the LabelAttributeName key that was used when the job was created in Ground Truth.
If your file is the output of a chained labeling job, specify the LabelAttributeName key for one or more jobs in the chain. Each LabelAttributeName key provides the annotations from an individual job.
Required: Yes
Type: Array of String
Minimum:
1Maximum:
63Update requires: Replacement
DocumentType-
The type of augmented manifest. PlainTextDocument or SemiStructuredDocument. If you don't specify, the default is PlainTextDocument.
-
PLAIN_TEXT_DOCUMENTA document type that represents any unicode text that is encoded in UTF-8. -
SEMI_STRUCTURED_DOCUMENTA document type with positional and structural context, like a PDF. For training with Amazon Comprehend, only PDFs are supported. For inference, Amazon Comprehend support PDFs, DOCX and TXT.
Required: No
Type: String
Allowed values:
PLAIN_TEXT_DOCUMENT | SEMI_STRUCTURED_DOCUMENTUpdate requires: Replacement
-
S3Uri-
The Amazon S3 location of the augmented manifest file.
Required: Yes
Type: String
Pattern:
^s3://[a-z0-9][\.\-a-z0-9]{1,61}[a-z0-9](/.*)?$Maximum:
1024Update requires: Replacement
SourceDocumentsS3Uri-
The S3 prefix to the source files (PDFs) that are referred to in the augmented manifest file.
Required: No
Type: String
Pattern:
^s3://[a-z0-9][\.\-a-z0-9]{1,61}[a-z0-9](/.*)?$Maximum:
1024Update requires: Replacement
Split-
The purpose of the data you've provided in the augmented manifest. You can either train or test this data. If you don't specify, the default is train.
TRAIN - all of the documents in the manifest will be used for training. If no test documents are provided, Amazon Comprehend will automatically reserve a portion of the training documents for testing.
TEST - all of the documents in the manifest will be used for testing.
Required: No
Type: String
Allowed values:
TRAIN | TESTUpdate requires: Replacement