对于 AI 代理:可在 https://www.mongodb.com/zh-cn/docs/llms.txt 获取文档索引—通过在任何 URL 路径后添加 .md 可获取所有页面的 Markdown 版本。
Docs 菜单

mongot 的推荐警报

此页面提供了一组用于自管理 mongot 部署的推荐 Prometheus 警报。警报定义是起始点。您可以复制、调整和优化这些警报定义以适应您的工作负载。

每个警报条目包含以下字段:

字段
说明

严重性

有关详情,请参阅 警报层级。

它告诉您什么

警报的操作含义。

PromQL

示例表达式。调整指标名称以适应您的环境。

阈值理由

给定阈值的原因。

第一响应

值班工程师应采取的操作。

首先为页面层级配置警报。运行这些警报一周,并调整警报的误报阈值。之后,添加工单和监视警报。

层级
何时发出警报

页面

客户数可见影响正在发生或即将发生。请尽快处理此警报。

工单

运行降级。在几小时内处理。

观看

在仪表盘上或用于趋势分析时很有用。无需立即操作。

以下警报表明客户可见影响,需要立即处理。

mongot 未响应 Prometheus 指标获取。搜索和向量搜索未正常运行。

使用以下 PromQL 表达式对此条件发出警报:

up{job="mongot"} == 0

将持续时间设置为一分钟。

短暂的失败可能是短暂的网络错误。持续超过一分钟的缺席是服务中断。

根据您运行 mongot 的方式,运行以下命令之一来响应此警报:

  • 对于使用 Kubernetes 的部署,运行 kubectl get pods。

  • 对于使用 systemd 的部署,运行 systemctl status mongot。

  • 对于使用 Docker 的部署,请运行 docker ps。

检查日志以了解崩溃原因。有关疑难排解步骤,请参阅 mongot 日志和 FTDC。

该进程反复重启。部署不稳定。

使用以下 PromQL 表达式对此条件发出警报:

changes(mongot_process_start_time_seconds[10m]) > 3

10 分钟内重启超过三次是崩溃循环。崩溃循环不是临时故障,需要处理。

通过执行以下操作来响应此警报:

  • 从最近的崩溃窗口捕获日志。

  • 暂停自动重启,以便检查已停止的 pod 或进程。

  • 打开 FTDC 捕获。

mongot 无法跟上 mongod。搜索结果会越来越过时。如果复制延迟未得到纠正,则游标会脱离oplog,并强制完全重新同步。

使用以下 PromQL 表达式之一来对此条件发出警报:

max(mongot_index_stats_indexing_replicationLagMs) > 60000

或者,要在达到绝对阈值之前捕获增长趋势,请使用:

deriv(max(mongot_index_stats_indexing_replicationLagMs)[15m:1m]) > 500

此指标按索引计算,单位为毫秒。请勿在 PromQL 中将指标除以 1000。稳态延迟低于一秒。一分钟的延迟在追赶场景中是可以接受的。持续增长的延迟是警报条件。

mongot_index_stats_* 指标系列仅在至少存在一个搜索索引时才会出现。在没有索引的全新部署中,此警报不会出现,因为该系列尚不存在。这是预期行为。

通过执行以下操作来响应此警报:

  • 检查 mongod 写入速率是否出现突发峰值。

  • 检查 mongot CPU 和磁盘 I/O 是否饱和。

有关指导,请参阅 mongot 的指标参考。

mongot 在同步过程中遇到错误。重复异常会强制重新同步。在重新同步过程中,索引暂时不可用或过时。

使用以下 PromQL 表达式之一来对此条件发出警报。

要根据同步过程中的错误发出警报,请使用:

increase(mongot_index_stats_indexing_steadyStateExceptions_total[10m]) > 0

要捕获初始同步异常,请使用:

increase(mongot_index_stats_indexing_initialSyncExceptions_total[10m]) > 0

生产中的任何稳态异常都是问题。oplog 已经滚动,这意味着 mongod oplog 太小或 mongot 太慢,或发生下游错误。

通过执行以下操作来响应此警报:

  • 捕获异常周围的 mongot 日志。

  • 检查 mongod oplog 大小。

  • 请立即打开FTDC捕获。事后确定根本原因难度更大。

Java虚拟机(JVM)堆几乎达到上限。即将发生 OutOfMemoryError。

使用以下 PromQL 表达式对此条件发出警报:

sum(mongot_jvm_memory_used_bytes{area="heap"})
/ sum(mongot_jvm_memory_max_bytes{area="heap"} > 0) > 0.85

将持续时间设置为五分钟。

持续堆占用率超过 85% 可能会导致问题。下一次分配峰值可能会导致内存不足错误。mongot 默认使用垃圾回收首先 (G1) 垃圾回收。以下表达式是上述 PromQL 表达式的更精确的后 GC 版本:

mongot_jvm_gc_live_data_size_bytes
/ mongot_jvm_gc_max_data_size_bytes > 0.85

通过检查活跃索引操作和查询负载来响应此警报。如果正在构建大型索引,则条件可能会在构建完成时解决。否则,增加 -Xmx 堆设置或减少并发操作的数量。

mongot 对 dataPath 卷上的磁盘使用情况实施三个阈值。这些阈值在 mongot 二进制本身中实施。无论您是否进行监控,阈值都会生效。在所有三个阈值处设置警报,以便值班工程师看到级联,并在超过最终阈值之前采取行动。

已用磁盘
mongot 的功能
对客户可见的影响
严重性

85% (15% 可用)

mongot 禁用初始同步。新索引构建保持在 PENDING 中。现有索引继续运行。

仅在创建新索引时可见。现有搜索和向量搜索功能照常运行。

工单

90% (10% 可用)

mongot 禁用稳态复制。现有索引停止从 mongod 接收更改事件。搜索结果越来越过时。

搜索结果与 mongod 过时。用户会看到最近写入的数据的过时结果。

页面

95% (5% 可用)

mongot 崩溃。恢复需要在 mongot 能够重新启动之前释放磁盘。

所有搜索和向量搜索均变为不可用。

页面

此警报有三条规则。严重性在每个阈值处逐渐升级。

使用以下 PromQL 表达式对这些条件发出警报:

85% - 工单级别:

(1 - mongot_system_disk_space_data_path_free_bytes
/ mongot_system_disk_space_data_path_total_bytes) >= 0.85

90% — 页面级别:

(1 - mongot_system_disk_space_data_path_free_bytes
/ mongot_system_disk_space_data_path_total_bytes) >= 0.90

95% — 页面级别(服务中断):

(1 - mongot_system_disk_space_data_path_free_bytes
/ mongot_system_disk_space_data_path_total_bytes) >= 0.95

mongot /metrics 终结点公开 _free_bytes 和 _total_bytes。将计算使用百分比设置为 1 - free/total。

通过执行以下操作来响应此警报:

  • 在 85% 处:通过搜索索引管理 API Atlas 审核并删除未使用的索引。切勿手动删除 dataPath 下的文件。新索引不会构建,直到磁盘使用率降至 85% 以下。

  • 当达到 90% 时:通知您的操作或 SRE 团队,集群处于复制禁用状态。删除未使用的索引或扩展存储以恢复复制。

  • 处于 95% 状态:这是一次服务中断。释放磁盘空间,然后重启 mongot。在释放磁盘之前,mongot 拒绝重新启动。

以下警报表明运行状况恶化,应在几小时内处理。

用户搜索速度变慢。

使用以下 PromQL 表达式之一来对此条件发出警报:

跨索引:

max(mongot_command_searchCommandTotalLatency_seconds{quantile="0.99"})
> <your-SLO-threshold>

每索引分解:

max(mongot_index_stats_query_searchResultBatchLatencies_seconds{quantile="0.99"})
by (indexId_logString) > <your-SLO-threshold>

阈值取决于您的服务级别目标。一个常见的起始点是:$search 的 99 百分位数在 500 毫秒内,$vectorSearch 的百分位数在一秒内。这些系列是带有预烘制 quantile 标签的摘要,而不是直方图。

通过调查以下内容来响应此警报:

  • 执行程序队列深度。

  • Java虚拟机(JVM)垃圾回收暂停时间。

  • 存储 IOPS 用于识别瓶颈。

工作人员已饱和,任务正在排队。查询延迟即将增加。

使用以下 PromQL 表达式对此条件发出警报:

max({__name__=~"mongot_.+_executor_queued_tasks"}) > 10

将持续时间设置为五分钟。

在负载峰值下,出现短暂的小队列是正常的。持续的队列意味着工作人员容量不足。常见的热点池包括:

  • mongot_decoding_executor

  • mongot_change_stream_sync_dispatcher_executor

  • mongot_indexing_work_executor

  • mongot_indexing_lifecycle_executor

  • mongot_index_commit_executor

索引工作分割到几个专用池中。没有组合的 mongot_indexing_executor。

通过确定哪个池正在排队来响应此警报:

topk(5, sum by (__name__) ({__name__=~"mongot_.+_executor_queued_tasks"}))

扩展或增加连接池大小。查询延迟升高之前,队列深度斜坡会提供早期警报。

存储卷正在接近饱和。Lucene 延迟越来越受磁盘限制。

使用以下 PromQL 表达式对此条件发出警报:

rate(mongot_system_disk_reads_events{name="<dataPath device>"}[5m]) > 1000

将时长设置为 15 分钟。

1、000 IOPS 阈值是存储类建议标志。但是,这不是一个严格的限制。正确的数字取决于您的设备。使用 df 或检查 mongot_system_disk_* 标签值来识别 dataPath 设备。

通过检查是否正在进行合并或初始同步来响应此警报。如果高 IOPS 级别持续存在,则存储类可能规格过小。重新检查存储配置。

操作系统会从磁盘重复拉取索引页面,因为它们已从缓存中驱逐。关键在于内存,而不是存储容量。

使用以下 PromQL 表达式对此条件发出警报:

rate(mongot_system_process_majorPageFaults_operations[5m]) > 1000

每秒 1,000 次主要缺页是内存压力处于关键路径上的规范阈值。结合持续的 IOPS,这就是相对于工作集而言内存不足的信号。

处理此警报:添加内存,因为 Lucene 会将索引文件内存映射。

特定索引遇到了非常见的索引失败。

使用以下 PromQL 表达式之一来对此条件发出警报:

increase(mongot_lifecycle_failedInitializationIndexes_total[10m]) > 0

或者:

increase(mongot_indexing_steadyStateChangeStream_unexpectedBatchFailures_total[10m]) > 0

或者:

increase(mongot_index_stats_indexing_invalidGeometryField_total[10m]) > 0

在正常负载下,这些计数器不会增加。增加表示存在数据问题,例如:

  • 映射爆炸。

  • 过大的文档。

  • 无效文档。

这些计数器的增加也可能表明代码路径问题。

通过执行以下操作来响应此警报:

  • 检查受影响索引的标签和原因。

  • 请检查 mongot 日志以获取根本的异常。

自动嵌入遇到问题。受影响的索引上的新文档的索引过程停滞。

使用以下 PromQL 表达式之一来对此条件发出警报:

increase(mongot_indexing_steadyStateChangeStream_rescheduledEmbeddingGetMores_total[10m]) > 0

或者:

increase(mongot_initialsync_queue_requeuedEmbeddingInitialSyncs_total[10m]) > 0

持续重新调度或重新入队表明嵌入路径未排干。最常见的原因是:

  • An invalid API key.

  • 不可访问的网络终结点。

  • Voyage AI 速率限制。

通过执行以下操作来响应此警报:

  • 检查 mongot 日志中针对嵌入终结点的 HTTP 错误。

  • 验证 API 密钥的有效性和连接性。

  • 检查 Voyage AI 状态。

一个或多个索引从 STEADY 状态过渡到恢复、过时或失败状态。

使用以下 PromQL 表达式对此条件发出警报:

count by (status) (mongot_index_stats_indexStatusCode{status!="STEADY"} == 1) > 0

在部署过程中,RECOVERING_TRANSIENT 中的单个索引在几秒内是正常的。以下任何状态中的持续计数大于零表示存在问题:

  • FAILED.

  • RECOVERING_NON_TRANSIENT.

  • STALE.

通过确定受影响的 indexId_logString 并检查对应的 mongot 日志行来响应此警报。

诊断捕获管道失败。mongot 在其他方面正常,但您已经失去了该节点的可观测性。

注意

此指标受 ftdcExecutorMetricsToPrometheus 功能标志控制。在添加此警报之前,请确认您的部署是否暴露此指标。此指标在默认 mongodb/mongodb-community-search 抓取中不存在。默认情况下,自管理部署的此标志处于关闭状态。

使用以下 PromQL 表达式对此条件发出警报:

mongot_mongot_ftdc_executor_failure_total > 0

将持续时间设置为五分钟。

如果您的部署暴露此指标,请将其视为下游可观测性降级的严重信号。

通过重启 mongot 响应此警报。

以下指标在仪表盘上对趋势分析很有用。这些指标都不需要分页。

此指标显示在垃圾回收后堆利用率的变化趋势。

使用以下 PromQL 表达式对此条件发出警报:

sum(mongot_jvm_gc_live_data_size_bytes) / sum(mongot_jvm_gc_max_data_size_bytes)

如果指标在几周内上升,请通过调查响应此警报。

此指标显示收集器中最近暂停的最差情况。

使用以下 PromQL 表达式对此条件发出警报:

max(mongot_jvm_gc_pause_seconds_max)

通过调查此指标是否在 100 毫秒内持续发出警报来响应此警报。

该指标显示开启文件描述符的软限制的内存余量。

使用以下 PromQL 表达式对此条件发出警报:

mongot_process_*

调查此指标是否超过 80% 来响应此警报。

此指标显示持有超时游标的客户端数量。

使用以下 PromQL 表达式对此条件发出警报:

rate(mongot_cursorManager_trackedCursors[5m])

无需为此指标设置警报阈值。纯粹信息性。

此指标显示等待连接的线程数。

使用以下 PromQL 表达式对此条件发出警报:

mongot_mongoClient_connectionPool_connectionsCheckedOut approaching _maxSize

如果指标持续存在且大于零,则通过调查响应此警报。

此指标显示存储容量。

使用以下 PromQL 表达式对此条件发出警报:

mongot_system_disk_space_data_path_free_bytes / mongot_system_disk_space_data_path_total_bytes

如果此指标低于 30% 可用,请考虑进行规划对话以增加存储。

Envoy exposes metrics for the availability of search and vector search, and for routing failures between mongod and mongot. Configure alerts for the following conditions. Use the equivalent metric names and labels exposed by your Envoy deployment.

Use the following metrics:

  • envoy_mongorpc_grpc_transcode_request_failed_total

  • envoy_mongorpc_grpc_transcode_request_total

  • mongot_index_stats_query_internallyFailedQueries_total

Calculate availability using the following:

1 - (Envoy failed requests + mongot internal failures) / Envoy total requests

Treat a sustained value below 0.70 as a high-severity alert.

Use envoy_mongorpc_grpc_transcode_request_failed_total and envoy_mongorpc_grpc_transcode_request_total. Alert when the following ratio is greater than 0.10 on at least 20% of active instances with more than 1 QPS, and at least 10 instances meet the QPS condition:

rate(envoy_mongorpc_grpc_transcode_request_failed_total[5m])
/ rate(envoy_mongorpc_grpc_transcode_request_total[5m]) > 0.10

Use envoy_cluster_upstream_rq_active{envoy_cluster_name="xds_cluster"} to identify Envoys with no active xDS upstream requests. Alert when fewer than 70% of Search Envoy instances are connected to the Search xDS server.

Use envoy_server_uptime. Treat envoy_server_uptime < 1800 as an instance-level signal and alert when more than 20% of Envoy instances meet the condition. Treat this as a supporting signal because small regions can produce false positives.

Use mongot_process_uptime_seconds. Alert when absent_over_time(mongot_process_uptime_seconds[15m]) is true for an instance that is expected to be running and scraped.

Use mongot_system_disk_space_data_path_free_bytes and mongot_system_disk_space_data_path_total_bytes. Alert when:

mongot_system_disk_space_data_path_free_bytes
/ mongot_system_disk_space_data_path_total_bytes < 0.10

Correlate this alert with replication pause, stale-index, and process-crash signals.

Use mongot_index_stats_indexing_replicationLagMs. Treat a maximum lag above 1800000 milliseconds (30 minutes) and rising as an early warning. Escalate above 7200000 milliseconds (2 hours) when the value continues to rise.

Use mongot_system_memory_phys_inUse_bytes, mongot_jvm_memory_used_bytes{area="heap"}, and mongot_system_memory_phys_total_bytes to calculate memory outside the JVM heap. Investigate when more than 5% of mongot instances exceed 50% native-memory usage for 55 minutes.

These thresholds are runbook starting points. Tune them for the size and workload of your deployment, and publish each alert only when the corresponding metric is available.

给本页内容打分