Overload occurs when your cluster receives more incoming operations than its resources can process, and it can cause a total or near-total outage on your cluster. Load Shedding, a feature of Intelligent Workload Management (IWM), allows Atlas to reject operations during prolonged overload. To configure Load Shedding for your cluster, see Configure Intelligent Workload Management.
When Load Shedding rejects an operation, your application might see a new error for that operation with the SystemOverloadedError label. This error indicates that your cluster is overloaded and shedding operations.
If you use a backpressure-aware client library, the client library recognizes these errors and retries the ones that are safe to retry. Your application still controls how long to keep retrying and when to shed load. To learn more, see Handle Overload Errors.
The SystemOverloadedError label on its own does not mean that the operation can be safely retried. To determine if an overload error is retryable, check for the following labels:
RetryableError- the operation was not executed and is safe to retryNoWritesPerformed- the server rejected the operation before performing any writes
The following example shows an overload error that Atlas returns when Load Shedding rejects an operation:
{ "ok": 0.0, "errmsg": "Request rejected: ingress operation rate limit exceeded", "code": 463, "codeName": "IngressOperationRateLimitExceeded", "errorLabels": ["SystemOverloadedError", "RetryableError", "NoWritesPerformed"] }
Important
If an overload error does not include the RetryableError label, ensure that the operation is idempotent and wait before retrying to avoid contributing to overload. Use exponential backoff and jitter in your retry logic.
Backpressure-Aware Client Libraries Versions
Backpressure-aware drivers and other client libraries automatically recognize overload errors with the SystemOverloadedError label and treat them as a signal of overload. If the error has a label that induces a retry, including the RetryableError label, the backpressure-aware client library automatically retries the operation with exponential backoff and jitter.
The following table lists the earliest client library versions that are backpressure-aware:
Client Library | Earliest Backpressure-Aware Version |
|---|---|
C Driver | 2.5 |
C++ Driver | 4.6 |
.NET/C# Driver | 3.11 |
Go Driver | 2.9 |
Java Sync Driver | 5.12 |
Java Reactive Streams Driver | 5.12 |
Kotlin Coroutine Driver | 5.12 |
Kotlin Sync Driver | 5.12 |
Node.js Driver | 7.6 |
PHP Library | 2.5; requires |
PyMongo | 4.18 |
Scala Driver | 5.12 |
Ruby | 2.26 |
Rust | 3.9 |
Handle Overload Errors
Handle overload errors in your application even if you use a backpressure-aware client library. Backpressure-aware client libraries retry operations labeled as retryable, but your application decides how long to keep retrying and when to shed load by abandoning operations that the cluster continues to reject. If you are not using a backpressure-aware client library, also implement the error detection and retry logic yourself.
See the following procedure for examples of how to implement error detection and retry logic with exponential backoff to handle overload errors that are labeled as retryable:
Use the following procedure to implement utilities to detect overload errors and retry with exponential backoff.
Guidelines for Safe Overload Retries
When retrying operations that failed due to an overload error, use the following guidelines to avoid contributing to overload and to increase the chances of successful retries:
Limit your retry attempts: Use a maximum of two retries per operation. Higher limits reduce error rates but increase server load during overload, while fewer retry attempts can reduce server load but increase errors rates.
Apply selectively: Use this pattern only for latency-sensitive or business-critical operations. For background workloads, log errors and retry at a higher level with longer delays.