WebCrawlerLimits¶
Structure Class¶
WebCrawlerLimits
dataclass
¶
The rate limits for the URLs that you want to crawl. You should be authorized to crawl the URLs.
Attributes¶
max_pages
class-attribute
instance-attribute
¶
max_pages: int | None = None
The max number of web pages crawled from your source URLs, up to 25,000 pages. If the web pages exceed this limit, the data source sync will fail and no web pages will be ingested.
rate_limit
class-attribute
instance-attribute
¶
rate_limit: int | None = None
The max rate at which pages are crawled, up to 300 per minute per host.