View a markdown version of this page

为亚马逊服务设置托管的 Prometheus 收集器 OpenSearch - Amazon Managed Service for Prometheus

本文属于机器翻译版本。若本译文内容与英语原文存在差异,则一律以英文原文为准。

为亚马逊服务设置托管的 Prometheus 收集器 OpenSearch

适用于 Prometheus 的亚马逊托管服务亚马逊 OpenSearch 服务的托管收集器会自动从您的 OpenSearch 服务域中抓取 Prometheus-compatible 指标并将其转发到您的目的地。适用于 Prometheus 的 Amazon Managed Service 为您管理收集器,为您提供所需的可扩展性、安全性和可靠性,无需自己管理任何实例、代理或抓取工具。您的目的地可以是 Prometheus 的亚马逊托管服务工作空间或亚马逊。 CloudWatch

您还可以创建与亚马逊弹性 Kubernetes 服务或适用于 Apache Kafka 的 Amazon Managed Streaming 集成的抓取工具。有关更多信息,请参阅集成亚马逊 EKS 和集成亚马逊 MSK 。

创建抓取程序

此过程假设您熟悉亚马逊 OpenSearch 服务域管理和亚马逊虚拟私有云网络概念。

为 OpenSearch 服务域创建抓取工具有几个先决条件:

  • 具有亚马逊虚拟私有云访问权限的亚马逊 OpenSearch 服务域。托管收集器仅支持具有 VPC 访问权限的 OpenSearch 服务域。不支持具有公有访问权限的域。

  • 允许托管收集器通过 HTTPS(端口 443)访问您的 OpenSearch服务域终端节点的安全组。向域的安全组添加入站规则,以允许来自为收集器提供的安全组的 HTTPS 流量。

要告诉抓取器要从哪个 OpenSearch 服务域中收集,请在请求exporters字段中指定域。该exporters字段获取导出器配置列表;对于 OpenSearch 服务域,请openSearchConfiguration提供该domainArn域的。您在source字段中单独提供网络(子网和安全组),如以下示例所示。

注意

抓取器使用创建抓取器时解析的域的 VPC 终端节点。如果您重新创建该域,它将获得一个新的端点,因此您必须创建一个新的抓取器或更新现有抓取器UpdateScraper才能从中收集。

创建抓取工具时需要提供抓取配置。托管式收集器会连接到指定的域,并自动收集其指标,因此无需在配置中指定抓取目标。配置必须包含 scrape_configs 部分,其中任务的 job_name 必须恰好为 opensearch-exporter。不包含名为的作业的配置opensearch-exporter将被拒绝。使用抓取配置来设置抓取间隔,也可以筛选或重新标记收集的指标。以下是抓取配置示例:

global: external_labels: domain_name: my-opensearch-domain scrape_configs: - job_name: opensearch-exporter scrape_interval: 60s

有关支持的抓取配置选项的更多信息,请参阅抓取程序配置。

To create a scraper using the AWS API

使用 CreateScraper API 操作通过 API 创建抓取 AWS 工具。以下示例在美国东部(弗吉尼亚北部)地区创建了一个抓取工具,该抓取器从 OpenSearch服务域收集指标并将其发送到适用于 Prometheus 的亚马逊托管服务工作区。用您自己的域名、网络和工作空间信息替换example内容,并提供您的抓取器配置。

注意

配置安全组和子网以匹配您的 OpenSearch 服务域的 Amazon VPC。在两个可用区中包括至少两个子网。

POST /scrapers HTTP/1.1 { "alias": "myScraper", "source": { "vpcConfiguration": { "securityGroupIds": ["sg-security-group-id"], "subnetIds": ["subnet-subnet-id-1", "subnet-subnet-id-2"] } }, "exporters": [ { "openSearchConfiguration": { "domainArn": "arn:aws:es:us-east-1:123456789012:domain/my-opensearch-domain" } } ], "destination": { "ampConfiguration": { "workspaceArn": "arn:aws:aps:us-east-1:123456789012:workspace/ws-workspace-id" } }, "scrapeConfiguration": { "configurationBlob": "base64-encoded-blob" } }

要将指标发送到亚马逊 CloudWatch 而不是 Prometheus 的亚马逊托管服务工作区,请将destination替换为配置: CloudWatch

"destination": { "cloudWatchConfiguration": { "datasetArn": "arn:aws:cloudwatch:us-east-1:123456789012:dataset/default" } }

该scrapeConfiguration参数需要一个 base64 编码的 Prometheus 配置 YAML 文件。运行以下命令之一来将 YAML 文件转换为 base64。也可以使用任何在线 base64 转换器来转换文件。

例 Linux/macOS
base64 -w0 scraper-config.yaml
例 Windows PowerShell
[Convert]::ToBase64String([System.IO.File]::ReadAllBytes("scraper-config.yaml"))
To create a scraper using the AWS CLI

通过 AWS Command Line Interface使用 create-scraper 命令创建抓取器。以下示例在美国东部(弗吉尼亚州北部)区域中创建抓取器。用您自己的域名、网络和工作空间信息替换example内容,并提供您的抓取器配置。

注意

配置安全组和子网以匹配您的 OpenSearch 服务域的 Amazon VPC。在两个可用区中包括至少两个子网。

aws amp create-scraper \ --source '{"vpcConfiguration":{"securityGroupIds":["sg-security-group-id"],"subnetIds":["subnet-subnet-id-1","subnet-subnet-id-2"]}}' \ --exporters '[{"openSearchConfiguration":{"domainArn":"arn:aws:es:us-east-1:123456789012:domain/my-opensearch-domain"}}]' \ --scrape-configuration configurationBlob=base64-encoded-blob \ --destination '{"ampConfiguration":{"workspaceArn":"arn:aws:aps:us-east-1:123456789012:workspace/ws-workspace-id"}}'
  • 以下是您可以在 AWS API 中使用的抓取器操作的完整列表:

    使用 CreateScraper API 操作创建抓取程序。

  • 使用 ListScrapers API 操作列出您现有的抓取程序。

  • 使用 UpdateScraper API 操作更新抓取工具的别名、配置或目的地。

  • 使用 DeleteScraper API 操作删除抓取程序。

  • 通过 DescribeScraper API 操作获取有关抓取程序的更多详细信息。

Cross-account 设置

要在跨账户设置中创建抓取工具,当您要从中收集指标的 OpenSearch 服务域与亚马逊 Prometheus 托管服务收集器位于不同的账户中时,请使用以下步骤。

例如,您有两个账户:一个是 OpenSearch 服务域account_id_source所在的源账户,另一个是亚马逊普罗米修斯管理服务工作空间account_id_target所在的目标账户。

注意

以下步骤中的信任策略通过其 ARN 引用抓取器,其中包括在您创建抓取器之前不scraper-id存在的。为避免这种排序问题,首先使用通配符 (scraper/*) 代替特定的抓取器 ARN 创建角色,创建抓取工具,然后更新两个信任策略,将通配符替换为返回的实际抓取器 ARN。CreateScraper

在跨账户设置中创建抓取器
  1. 在源账户中,创建角色 arn:aws:iam::111122223333:role/Source 并添加以下信任策略。

    { "Effect": "Allow", "Principal": { "Service": [ "scraper.aps.amazonaws.com" ] }, "Action": "sts:AssumeRole", "Condition": { "ArnEquals": { "aws:SourceArn": "arn:aws:aps:aws-region:111122223333:scraper/scraper-id" }, "StringEquals": { "AWS:SourceAccount": "111122223333" } } }
  2. 在目标账户中,创建一个角色arn:aws:iam::444455556666:role/Target并添加以下信任策略,允许源角色代入该角色。

    { "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam::111122223333:role/Source" }, "Action": "sts:AssumeRole", "Condition": { "StringEquals": { "sts:ExternalId": "arn:aws:aps:aws-region:111122223333:scraper/scraper-id" } } }

    为目标角色附加权限策略,允许其写入您的目标。如果您的目的地是 Prometheus 的亚马逊托管服务工作空间,请附加 AmazonPrometheusRemoteWriteAccess(即授权)。aps:RemoteWrite如果您的目的地是亚马逊 CloudWatch,请附cloudwatch:PutMetricData上对数据集授予的政策。

  3. 使用 --role-configuration 选项创建抓取器。

    aws amp create-scraper \ --source '{"vpcConfiguration":{"securityGroupIds":["sg-security-group-id"],"subnetIds":["subnet-subnet-id-1","subnet-subnet-id-2"]}}' \ --exporters '[{"openSearchConfiguration":{"domainArn":"arn:aws:es:aws-region:111122223333:domain/my-opensearch-domain"}}]' \ --scrape-configuration configurationBlob=<base64-encoded-blob> \ --destination '{"ampConfiguration":{"workspaceArn":"arn:aws:aps:aws-region:444455556666:workspace/ws-workspace-id"}}' \ --role-configuration '{"sourceRoleArn":"arn:aws:iam::111122223333:role/Source", "targetRoleArn":"arn:aws:iam::444455556666:role/Target"}'
  4. 验证抓取器创建。

    aws amp list-scrapers

查找和删除抓取程序

您可以使用 AWS API 或 AWS CLI 来列出账户中的抓取工具或将其删除。

注意

确保您使用的是 AWS CLI 或 SDK 的最新版本。最新版本为您提供最新的特征和功能,以及安全更新。或者使用 AWS CloudShell,它能自动提供始终最新的命令行体验。

要列出您账户中的所有抓取程序,请使用 ListScrapers API 操作。或者,使用 AWS CLI,致电:

aws amp list-scrapers

要删除抓取程序,请使用 ListScrapers 操作查找要删除的抓取程序的 scraperId,然后使用 DeleteScraper 操作将其删除。或者,使用 AWS CLI,致电:

aws amp delete-scraper --scraper-id scraperId

从亚马逊 OpenSearch 服务收集的指标

当您与亚马逊 OpenSearch 服务集成时,亚马逊 Prometheus 托管服务收集器会自动抓取描述您的域名运行状况和 Prometheus-compatible 性能的指标。收集器发出数百个指标;下表列出了每个类别的代表性指标。下表中的opensearch_indices_指标是跨域汇总的,许多指标都有仅限主变体和总变体(后缀_primary为 or)。_total收集器还会为每个索引、前缀和每分片指标发出相同的统计数据opensearch_index_stats_,前缀为前缀。opensearch_indices_shards_

指标 描述/用途

打开搜索集群运行状况状态

集群运行状况(绿色、黄色或红色)。

opensearch_cluster_health_health_node_

群集中节点的数量。

打开搜索集群运行状况数据节点数

集群中数据节点的数量。

opensearch_cluster_health_active_primary_shards

活跃的主分片数量。

opensearch_cluster_health_active_shards

活动分片总数,包括主分片和副本分片。

opensearch_cluster_health_relocating_shards

在节点之间重新定位的分片数量。

opensearch_cluster_health_health_初始化分片

正在初始化的分片数量。

opensearch_cluster_health_unsigned_shards

未分配给节点的分片数量。

opensearch_cluster_health_delayed_unsigned_shards

分配延迟的未分配分片的数量。

打开搜索集群运行状况待处理任务数

排队等待运行的集群级任务的数量。

opensearch_cluster_health_in_flight_fetch 的生命值

正在进行的分片提取请求的数量。

opensearch_cluster_health_task_max_waiting_in_queue_millis

任务在队列中等待的最长时间,以毫秒为单位。

指标 描述/用途

opensearch_os_cpu_percent

节点上操作系统使用的 CPU 百分比。

opensearch_os_load1

过去 1 分钟的操作系统平均负载。收集器还会发射opensearch_os_load5和。opensearch_os_load15

opensearch_os_mem_used_bytes

使用的物理内存量,以字节为单位。收集器还发出opensearch_os_mem_free_bytesopensearch_os_mem_actual_used_bytes、和。opensearch_os_mem_actual_free_bytes

opensearch_jvm_memory_used_bytes

按内存区域(堆和非堆)划分的 JVM 内存使用量,以字节为单位。相关指标包括 opensearch_jvm_memory_committed_bytes 和 opensearch_jvm_memory_max_bytes。

opensearch_jvm_gc_collection_seconds_count

JVM 垃圾回收事件的总数。与垃圾收集所花费的总时间一起opensearch_jvm_gc_collection_seconds_sum使用。

opensearch_jvm_uptime_seconds

JVM 正常运行时间,以秒为单位。

opensearch_thread_pool_active_count

每个线程池中的活动线程数。相关指标包括opensearch_thread_pool_queue_countopensearch_thread_pool_rejected_count、和opensearch_thread_pool_completed_count。

opensearch_filesystem_data_ailable_bytes

节点可用的磁盘空间量,以字节为单位。相关指标包括 opensearch_filesystem_data_free_bytes 和 opensearch_filesystem_data_size_bytes。

opensearch_filesystem_io_stats_device_read_operations_count

磁盘读取操作的次数。收集器还会发出写入和总操作次数以及读取和写入大小(例如,opensearch_filesystem_io_stats_device_read_size_kilobytes_sum)。

opensearch_process_cpu_per

OpenSearch进程使用的 CPU 百分比。

opensearch_process_mem_resident_size_by

OpenSearch 进程的常驻内存大小,以字节为单位。

打开搜索进程打开文件数

该 OpenSearch进程打开的文件描述符的数量。

opensearch_breakers_triped

每个断路器跳闸的总次数。相关指标包括 opensearch_breakers_estimated_size_bytes 和 opensearch_breakers_limit_size_bytes。

打开搜索_transport_rx_size_bytes_total

通过节点间传输层接收的数据总量,以字节为单位。收集器还会发射opensearch_transport_tx_size_bytes_total数据包数。

opensearch_indexing_pressure_current_all_in_bytes

索引请求消耗的当前内存,以字节为单位,opensearch_indexing_pressure_limit_in_bytes以配置的限制为准。

指标 描述/用途

打开搜索索引文档

文件数量。相关指标包括opensearch_indices_docs_deletedopensearch_indices_docs_primary、和opensearch_indices_docs_total。

打开搜索索引存储大小字节

磁盘上索引的总大小,以字节为单位,包含opensearch_indices_store_size_bytes_primary和opensearch_indices_store_size_bytes_total变体。

打开搜索索引_索引_索引_总计

已编入索引的文档总数。与一起使用可opensearch_indices_indexing_index_time_seconds_total实现索引延迟。

opensearch_indexing_indexing_is_throttled

表示索引当前是否受到限制,包括opensearch_indices_indexing_throttle_time_seconds_total总限制时间。

opensearch_indices_search_query_total

搜索查询的总数。与一起使用opensearch_indices_search_query_time_seconds可延长查询延迟。

opensearch_indices_search_fetch_total

提取操作总数,包括opensearch_indices_search_fetch_time_seconds提取延迟。

opensearch_indices_get_total

获取操作的总数,包括opensearch_indices_get_time_seconds获取延迟。

opensearch_indices_merges_total

已完成的区段合并总数。相关指标包括 opensearch_indices_merges_current 和 opensearch_indices_merges_total_time_seconds_total。

打开搜索索引刷新总数

索引刷新操作总数,包括opensearch_indices_refresh_time_seconds_total刷新时间。

打开搜索索引_flush_total

索引刷新操作总数,包括opensearch_indices_flush_time_seconds刷新时间。

打开搜索索引区段数

分段数。收集器还会发出每段内存指标(例如,opensearch_indices_segments_memory_bytes和opensearch_indices_segment_terms_memory_total)。

打开搜索索引_translog_操作

事务日志中的操作数及其大小。opensearch_indices_translog_size_in_bytes

opensearch_indices_field 数据_内存_大小_字节

字段数据缓存使用的内存,以字节为单位,opensearch_indices_fielddata_evictions用于驱逐。

opensearch_indices_query_cache_memory_size_bytes

查询缓存使用的内存,以字节为单位。相关指标包括 opensearch_indices_query_cache_count 和 opensearch_indices_query_cache_evictions。

打开搜索索引请求缓存内存大小字节

请求缓存使用的内存,以字节为单位。相关指标包括 opensearch_indices_request_cache_count 和 opensearch_indices_request_cache_evictions。

限制

亚马逊 OpenSearch 服务与亚马逊 Prometheus 托管服务的集成存在以下限制:

  • 仅支持具有亚马逊虚拟私有云访问权限的 OpenSearch 服务域。不支持具有公有访问权限的域。

  • 抓取器从单个 OpenSearch 服务域收集指标。要从多个域收集指标,请为每个域创建单独的抓取程序。