此页面介绍如何将 mongot 指标和日志与常用监控平台集成。这些说明假设您已经运行了其中一个工具,并且需要 mongot 特定配置。此页面不会从头开始教您使用 Prometheus、Grafana 或其他平台。
本指导针对将 mongot 集成到现有可观测性堆栈中的站点可靠性工程师和平台团队。
可用的监控界面
下表汇总了 mongot 曝露的用于监控的界面:
表面 | protocol | 默认终结点 | 在以下配置 | 注意 |
|---|---|---|---|---|
衡量标准 | HTTP, Prometheus text format |
|
| 包含的 |
活跃性 | HTTP |
|
| 在 |
就绪程度 | HTTP |
|
| 当 |
日志 | stdout 和 stderr,或文件 | 每个日志记录配置 |
| 根据配置,为 JSON 或文本。 |
FTDC | 磁盘二进制流 |
|
| 默认启用。使用 |
Prometheus 和 Grafana
Prometheus 和 Grafana 是大多数自管理部署的推荐监控堆栈。该堆栈免费、受到广泛支持,并可直接与 mongot 指标终结点配合使用。Prometheus 可在 Community 和 Enterprise 版本上运行,不需要 Ops Manager 配置。
抓取配置
将抓取作业添加到 Prometheus 配置中:
scrape_configs: - job_name: mongot scrape_interval: 15s scrape_timeout: 10s static_configs: - targets: - mongot-host-1.internal:9946 - mongot-host-2.internal:9946 labels: deployment: prod edition: ce
对于 Kubernetes Operator 所管理的 Kubernetes 部署,请使用带有 Prometheus Operator 的 PodMonitor 或 ServiceMonitor 资源。定位标记为 app=<resource-name>-search:的 Pod
apiVersion: monitoring.coreos.com/v1 kind: PodMonitor metadata: name: mongot namespace: <mongot-namespace> spec: selector: matchLabels: app: <resource-name>-search podMetricsEndpoints: - port: metrics interval: 15s
记录规则
记录规则可减少 PromQL 重复,并使 Grafana 查询更快。以下规则使用自管理 mongot 公开的指标名称:
groups: - name: mongot_recording interval: 30s rules: - record: mongot:search_latency_p99 expr: max(mongot_command_searchCommandTotalLatency_seconds{quantile="0.99"}) - record: mongot:vector_search_latency_p99 expr: max(mongot_command_vectorSearchCommandTotalLatency_seconds{quantile="0.99"}) - record: mongot:search_rate:rate5m expr: sum(rate(mongot_command_searchCommandTotalLatency_seconds_count[5m])) - record: mongot:search_failure_rate:rate5m expr: sum(rate(mongot_command_searchCommandFailure_total[5m])) - record: mongot:replication_lag_ms:max expr: max(mongot_index_stats_indexing_replicationLagMs) - record: mongot:heap_utilization_post_gc expr: mongot_jvm_gc_live_data_size_bytes / mongot_jvm_gc_max_data_size_bytes - record: mongot:gc_pause_worst expr: max(mongot_jvm_gc_pause_seconds_max)
警报规则
将推荐的警报转换为 Prometheus 警报规则。示例:
groups: - name: mongot_alerts rules: - alert: MongotDown expr: up{job="mongot"} == 0 for: 1m labels: severity: page annotations: summary: "mongot is down ({{ $labels.instance }})" - alert: MongotReplicationLagGrowing expr: deriv(max(mongot_index_stats_indexing_replicationLagMs)[15m:1m]) > 500 for: 10m labels: severity: page - alert: MongotHeapPressure expr: mongot:heap_utilization_post_gc > 0.85 for: 5m labels: severity: page
要将全部警报翻译成 PromQL,请参阅 建议用于 mongot 的警报。
Grafana 仪表盘骨架
mongot 的初始 Grafana 仪表盘应包括以下面板组:
进程:运行时间、重启计数、CPU 和常驻内存。
Java虚拟机(JVM):堆内存使用情况与最大值的比较、GC 后堆内存、GC 暂停时间和线程。
复制:当前状态、延迟毫秒数和速率,以及每秒应用的事件。
索引:活动构建、每个索引的状态、索引失败和合并积压。
查询:按操作符分类的速率、按操作符分类的延迟 p50、p95 和 p99,以及错误率。
执行程序:按池和拒绝任务排队深度。
存储:可用字节、IOPS 和页面错误率。
嵌入(如果已启用):请求速率、延迟、错误和令牌吞吐量。
OpenTelemetry
如果您的组织对 OpenTelemetry 进行标准化,OpenTelemetry Collector 可以从 Prometheus 终结点捕获 mongot 指标并将其转发到任何 OTLP 兼容后端:
receivers: prometheus: config: scrape_configs: - job_name: mongot scrape_interval: 15s static_configs: - targets: - localhost:9946 exporters: otlphttp: endpoint: https://<your-otel-backend> service: pipelines: metrics: receivers: - prometheus exporters: - otlphttp
此模式与提供商无关。相同的采集器配置适用于 Honeycomb、Grafana Cloud、New Relic 和其他后端。
要转发日志,请配置mongot以将JSON写入stdout。收集器可以解析结构化字段并将日志路由到后端。
日志转发
mongot 默认将结构化 JSON 日志写入 stdout 和 stderr。将 stdout 转发到集中式日志平台,并将日志作为 JSON 吸收。
Fluent Bit 和 Vector
Fluent Bit 和 Vector 都可用于 mongot 日志采集。将日志视为标签流。要学习哪些日志模式最重要,请参阅mongot 日志和 FTDC。
Cloudwatch 日志
对于 AWS 托管部署,CloudWatch 代理可直接追踪日志文件。在关键日志模式(例如 Exception requiring resync)上创建 CloudWatch 指标过滤器,以将日志事件转换为指标。
健康检查
mongot 默认在端口 8080 上暴露两个 HTTP 终结点:
端点 | 用于 | 含义 |
|---|---|---|
| 活跃性 |
|
| 就绪程度 |
|
两个终结点都返回带有HTTP 200的JSON。将{"status":"SERVING"}视为健康,将{"status":"NOT_SERVING"}视为不健康。无效的查询参数返回带有{"error":"BAD_REQUEST"}的HTTP 400。
与指标终结点一样,/health 和 /ready 终结点默认不进行身份验证。在网络层保护它们。
在 Kubernetes 中,将活蹦探测映射到 /health,将就绪探测映射到 /ready:
livenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 30 periodSeconds: 10 failureThreshold: 3 readinessProbe: httpGet: path: /ready port: 8080 initialDelaySeconds: 5 periodSeconds: 5 failureThreshold: 3
如果对两个探针都使用 /health,则在其索引初始化之前,节点可以接收流量,因为服务绑定后,/health 会立即返回 SERVING。存在两个终结点分割是为了分离这些信号。
当 Kubernetes 操作符的 MongoDB 控制器管理多个 mongot 节点时,它会预配一个默认负载均衡器,并根据 /ready 终结点为您路由流量。对于在多个 mongot 实例前运行自己的负载均衡器的自管理部署,请将负载均衡器配置为仅向在 /ready 上返回 SERVING 的实例路由流量。
为了使 Pod 即使某些索引无法初始化也能保持就绪,请将就绪探测路径设置为 /ready?allowFailedIndexes=true。此设置是一种有意的权衡,因为失败的索引会对访问它们的查询返回空结果。
多实例注意事项
如果您的部署运行的 mongot 实例超过一个,则每个实例都会暴露其自己的指标终结点。分别刮取每个实例,然后使用 Prometheus 聚合(例如 sum、max 和 avg)来查看指标的组合视图。
追踪每个实例和总体的复制延迟、执行程序队列深度和查询延迟。单个饱和实例会降低其提供服务的查询延迟,而整个集群的平均值可能会隐藏这种降级。
对于分片集群,请为每个抓取标记分片名称,以便您可以按分片汇总指标。
FTDC 和支持案例
当您打开 MongoDB 支持工单时,请附上受影响的 mongot 实例的 FTDC 捕获。要了解捕获过程,请参阅 mongot 日志和 FTDC。