Set up managed Prometheus collectors for Amazon OpenSearch Service
The Amazon Managed Service for Prometheus managed collector for Amazon OpenSearch Service automatically scrapes Prometheus-compatible metrics from your OpenSearch Service domain and forwards them to your destination. Amazon Managed Service for Prometheus manages the collector for you, giving you the scalability, security, and reliability that you need, without having to manage any instances, agents, or scrapers yourself. Your destination can be an Amazon Managed Service for Prometheus workspace or Amazon CloudWatch.
You can also create a scraper that integrates with Amazon Elastic Kubernetes Service or with Amazon Managed Streaming for Apache Kafka. For more information, see Integrate Amazon EKS and Integrate Amazon MSK.
Create a scraper
This procedure assumes that you are familiar with Amazon OpenSearch Service domain administration and Amazon Virtual Private Cloud networking concepts.
There are a few prerequisites for creating a scraper for an OpenSearch Service domain:
-
An Amazon OpenSearch Service domain with Amazon Virtual Private Cloud access. Managed collectors support only OpenSearch Service domains that have VPC access. Domains with public access are not supported.
-
Security groups that allow the managed collector to reach your OpenSearch Service domain endpoint over HTTPS (port 443). Add an inbound rule to the domain's security group that allows HTTPS traffic from the security group that you provide for the collector.
To tell the scraper which OpenSearch Service domain to collect from, you specify
the domain in the exporters field of the request. The
exporters field takes a list of exporter configurations; for an
OpenSearch Service domain, provide an openSearchConfiguration
with the domainArn of the domain. You provide the networking (subnets
and security group) separately in the source field, as shown in the
following examples.
Note
The scraper uses the domain's VPC endpoint as resolved when you create the scraper. If you recreate the domain, it gets a new endpoint, so you must create a new scraper or update the existing one with UpdateScraper to collect from it.
You provide a scrape configuration when you create the scraper. The managed
collector connects to the domain that you specify and collects its metrics
automatically, so you do not specify scrape targets in the configuration. The
configuration must include a scrape_configs section with a job whose
job_name is exactly opensearch-exporter. A
configuration that does not contain a job named opensearch-exporter is
rejected. Use the scrape configuration to set the scrape interval and, optionally,
to filter or relabel the collected metrics. The following is an example scrape
configuration:
global: external_labels: domain_name:my-opensearch-domainscrape_configs: - job_name: opensearch-exporter scrape_interval: 60s
For more information about the supported scrape configuration options, see Scraper configuration.
-
The following is a full list of the scraper operations that you can use with the AWS API:
Create a scraper with the CreateScraper API operation.
-
List your existing scrapers with the ListScrapers API operation.
-
Update the alias, configuration, or destination of a scraper with the UpdateScraper API operation.
-
Delete a scraper with the DeleteScraper API operation.
-
Get more details about a scraper with the DescribeScraper API operation.
Cross-account setup
To create a scraper in a cross-account setup, when the OpenSearch Service domain that you want to collect metrics from is in a different account from the Amazon Managed Service for Prometheus collector, use the following procedure.
For example, you have two accounts: a source account
account_id_source where the OpenSearch Service domain is located,
and a target account account_id_target where the Amazon Managed Service for Prometheus workspace
resides.
Note
The trust policies in the following steps refer to the scraper by its ARN,
which includes a scraper-id that does not exist until
you create the scraper. To avoid this ordering problem, first create the roles
with a wildcard (scraper/*) in place of the specific scraper ARN,
create the scraper, and then update both trust policies to replace the wildcard
with the actual scraper ARN that CreateScraper returns.
To create a scraper in a cross-account setup
-
In the source account, create a role
arn:aws:iam::and add the following trust policy.111122223333:role/Source{ "Effect": "Allow", "Principal": { "Service": [ "scraper.aps.amazonaws.com" ] }, "Action": "sts:AssumeRole", "Condition": { "ArnEquals": { "aws:SourceArn": "arn:aws:aps:aws-region:111122223333:scraper/scraper-id" }, "StringEquals": { "AWS:SourceAccount": "111122223333" } } } -
In the target account, create a role
arn:aws:iam::and add the following trust policy, which allows the source role to assume it.444455556666:role/Target{ "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::111122223333:role/Source" }, "Action": "sts:AssumeRole", "Condition": { "StringEquals": { "sts:ExternalId": "arn:aws:aps:aws-region:111122223333:scraper/scraper-id" } } }Attach a permissions policy to the target role that allows it to write to your destination. If your destination is an Amazon Managed Service for Prometheus workspace, attach AmazonPrometheusRemoteWriteAccess (which grants
aps:RemoteWrite). If your destination is Amazon CloudWatch, attach a policy that grantscloudwatch:PutMetricDataon the dataset. -
Create a scraper with the
--role-configurationoption.aws amp create-scraper \ --source '{"vpcConfiguration":{"securityGroupIds":["sg-security-group-id"],"subnetIds":["subnet-subnet-id-1","subnet-subnet-id-2"]}}' \ --exporters '[{"openSearchConfiguration":{"domainArn":"arn:aws:es:aws-region:111122223333:domain/my-opensearch-domain"}}]' \ --scrape-configuration configurationBlob=<base64-encoded-blob>\ --destination '{"ampConfiguration":{"workspaceArn":"arn:aws:aps:aws-region:444455556666:workspace/ws-workspace-id"}}' \ --role-configuration '{"sourceRoleArn":"arn:aws:iam::111122223333:role/Source", "targetRoleArn":"arn:aws:iam::444455556666:role/Target"}' -
Validate the scraper creation.
aws amp list-scrapers
Find and delete scrapers
You can use the AWS API or the AWS CLI to list the scrapers in your account or to delete them.
Note
Make sure that you are using the latest version of the AWS CLI or SDK. The latest version provides you with the latest features and functionality, as well as security updates. Alternatively, use AWS CloudShell, which provides an always up-to-date command line experience, automatically.
To list all the scrapers in your account, use the ListScrapers API operation. Alternatively, with the AWS CLI, call:
aws amp list-scrapers
To delete a scraper, find the scraperId for the scraper that you want
to delete, using the ListScrapers operation, and then use the DeleteScraper operation to delete it. Alternatively, with the AWS CLI,
call:
aws amp delete-scraper --scraper-idscraperId
Metrics collected from Amazon OpenSearch Service
When you integrate with Amazon OpenSearch Service, the Amazon Managed Service for Prometheus collector
automatically scrapes Prometheus-compatible metrics that describe the health and
performance of your domain. The collector emits several hundred metrics; the
following tables list representative metrics from each category. The
opensearch_indices_ metrics in the following tables are aggregated
across the domain, and many have primary-only and total variants (suffixed
_primary or _total). The collector also emits the same
statistics per index, prefixed opensearch_index_stats_, and per-shard
metrics prefixed opensearch_indices_shards_.
| Metric | Description / Purpose |
|---|---|
opensearch_cluster_health_status |
Cluster health status (green, yellow, or red). |
opensearch_cluster_health_number_of_nodes |
Number of nodes in the cluster. |
opensearch_cluster_health_number_of_data_nodes |
Number of data nodes in the cluster. |
opensearch_cluster_health_active_primary_shards |
Number of active primary shards. |
opensearch_cluster_health_active_shards |
Total number of active shards, including primary and replica shards. |
opensearch_cluster_health_relocating_shards |
Number of shards that are relocating between nodes. |
opensearch_cluster_health_initializing_shards |
Number of shards that are initializing. |
opensearch_cluster_health_unassigned_shards |
Number of shards that are not assigned to a node. |
opensearch_cluster_health_delayed_unassigned_shards |
Number of unassigned shards whose assignment is delayed. |
opensearch_cluster_health_number_of_pending_tasks |
Number of cluster-level tasks that are queued and waiting to run. |
opensearch_cluster_health_number_of_in_flight_fetch |
Number of shard fetch requests in progress. |
opensearch_cluster_health_task_max_waiting_in_queue_millis |
Longest time that a task has waited in the queue, in milliseconds. |
| Metric | Description / Purpose |
|---|---|
opensearch_os_cpu_percent |
Percentage of CPU used by the operating system on the node. |
opensearch_os_load1 |
Operating system load average over the last 1
minute. The collector also emits
|
opensearch_os_mem_used_bytes |
Amount of physical memory used, in bytes. The
collector also emits
|
opensearch_jvm_memory_used_bytes |
Amount of JVM memory used, in bytes, by memory
area (heap and non-heap). Related metrics include
|
opensearch_jvm_gc_collection_seconds_count |
Total number of JVM garbage collection events. Use
with |
opensearch_jvm_uptime_seconds |
JVM uptime, in seconds. |
opensearch_thread_pool_active_count |
Number of active threads in each thread pool.
Related metrics include
|
opensearch_filesystem_data_available_bytes |
Amount of disk space available to the node, in
bytes. Related metrics include
|
opensearch_filesystem_io_stats_device_read_operations_count |
Number of disk read operations. The collector also
emits write and total operation counts and read and write
sizes (for example,
|
opensearch_process_cpu_percent |
Percentage of CPU used by the OpenSearch process. |
opensearch_process_mem_resident_size_bytes |
Resident memory size of the OpenSearch process, in bytes. |
opensearch_process_open_files_count |
Number of file descriptors open by the OpenSearch process. |
opensearch_breakers_tripped |
Total number of times each circuit breaker has
tripped. Related metrics include
|
opensearch_transport_rx_size_bytes_total |
Total amount of data received over the transport
layer between nodes, in bytes. The collector also emits
|
opensearch_indexing_pressure_current_all_in_bytes |
Current memory consumed by indexing requests, in
bytes, with |
| Metric | Description / Purpose |
|---|---|
opensearch_indices_docs |
Number of documents. Related metrics include
|
opensearch_indices_store_size_bytes |
Total size of the indices on disk, in bytes, with
|
opensearch_indices_indexing_index_total |
Total number of documents indexed. Use with
|
opensearch_indices_indexing_is_throttled |
Indicates whether indexing is currently throttled,
with |
opensearch_indices_search_query_total |
Total number of search queries. Use with
|
opensearch_indices_search_fetch_total |
Total number of fetch operations, with
|
opensearch_indices_get_total |
Total number of get operations, with
|
opensearch_indices_merges_total |
Total number of segment merges completed. Related
metrics include |
opensearch_indices_refresh_total |
Total number of index refresh operations, with
|
opensearch_indices_flush_total |
Total number of index flush operations, with
|
opensearch_indices_segments_count |
Number of segments. The collector also emits
per-segment memory metrics (for example,
|
opensearch_indices_translog_operations |
Number of operations in the transaction log, with
|
opensearch_indices_fielddata_memory_size_bytes |
Memory used by the field data cache, in bytes, with
|
opensearch_indices_query_cache_memory_size_bytes |
Memory used by the query cache, in bytes. Related
metrics include
|
opensearch_indices_request_cache_memory_size_bytes |
Memory used by the request cache, in bytes. Related
metrics include
|
Limitations
The Amazon OpenSearch Service integration with Amazon Managed Service for Prometheus has the following limitations:
-
Only supported for OpenSearch Service domains that have Amazon Virtual Private Cloud access. Domains with public access are not supported.
-
A scraper collects metrics from a single OpenSearch Service domain. To collect metrics from more than one domain, create a separate scraper for each domain.