For AI agents: a documentation index is available at https://www.mongodb.com/docs/llms.txt — markdown versions of all pages are available by appending .md to any URL path.
Docs Menu

Intelligent Workload Management

Note

This page applies to both Atlas Infinite and Atlas Core.

Modern applications must stay available even when traffic spikes and expensive queries consume more resources than expected. Intelligent Workload Management (IWM) is the MongoDB Atlas-managed availability capability that keeps your clusters responsive and predictable, even under overload.

Overload occurs when your cluster receives more operations than its resources can process. Overload can develop faster than your infrastructure can scale. Even if you enable auto-scaling, Atlas requires time to provision additional capacity. This opens a window during which a cluster can become unstable if it treats all operations the same way. IWM closes that window by monitoring and reacting to workload spikes and resource pressure in real time.

IWM is the overload protection system that Atlas applies to all Dedicated clusters running MongoDB 9.0 or later. IWM includes the following protections:

  • Always-on protections queue, slow down, and prioritize operations to prevent them from overwhelming the cluster. These protections form a baseline of defense against overload. To learn more, see Always-On Protections.

  • Load Shedding is an optional protection that rejects incoming database operations in response to sustained overload. To learn more, see Load Shedding.

Atlas applies a subset of these protections to each cluster, based on the cluster's tier, topology, MongoDB version, and Load Shedding configuration. To learn more, see Cluster and Version Requirements.

Always-on protections are the baseline set of IWM protections that Atlas applies to absorb overload and maintain availability. These protections run at all times, even when Load Shedding is disabled, and you can't turn them off. To learn which of these protections apply to your cluster, see Cluster and Version Requirements.

  • Ingress queueing: Atlas runs an ingress request rate limiter on each mongod process in your cluster, which determines how quickly each process admits operations for execution. When a traffic spike outlasts your cluster's reserved capacity, the rate limiter adds excess incoming operations to an ingress queue instead of letting them compete for resources with running operations. If resource use stays high, the rate limiter admits operations at a slower rate, so more operations wait in the queue. The rate limiter rejects operations only if you enable Load Shedding on your cluster.

    If reactive auto-scaling is enabled on your cluster, Atlas uses the rate of queued operations to determine when to scale your cluster up. Atlas reports this rate as Operation Rate Limiting: Queued operations in Atlas metrics. To learn more, see How Atlas Scales Cluster Tier.

  • Execution control prioritization: Atlas runs critical reads and writes and internal cluster activity, such as elections, ahead of background jobs and expensive queries. Under pressure, Atlas slows expensive, long-running, or low-priority operations first.

  • Cache eviction and memory tuning: Atlas tunes how your cluster uses memory and evicts cached data. When memory runs low, your cluster slows down in a controlled way instead of running out of memory or repeatedly clearing the data it needs most.

  • Connection rate limiting: During connection storms, Atlas admits new client connections at a safe rate and rejects excess connection requests while preserving existing connections. To learn more, see Connection Rate Limits.

  • Write-blocking: On Atlas Core clusters, if the primary node runs critically low on free disk space, Atlas blocks write operations on the replica set. Atlas continues to serve reads while writes are blocked. Blocked write operations include delete operations, because deleting documents can temporarily increase used disk space. On sharded clusters, Atlas blocks only writes routed to the affected shard.

    A node's total disk size determines the thresholds Atlas applies. The following table lists the thresholds for each disk size range:

    Total Disk Size
    Thresholds

    8 GiB to 20 GiB

    Atlas blocks writes when free disk space drops below 600 MiB. Atlas unblocks writes when free disk space exceeds 900 MiB.

    20 GiB to 1.25 TiB

    Atlas blocks writes when free disk space drops below 4% of total disk size. Atlas unblocks writes when free disk space exceeds 6% of total disk size.

    1.25 TiB or greater

    Atlas blocks writes when free disk space drops below 50 GiB. Atlas unblocks writes when free disk space exceeds 75 GiB.

    To recover disk space, increase cluster storage or enable storage auto-scaling.

    When Atlas blocks writes, Atlas logs return a MongoServerError indicating that writes are blocked. The error code and reason depend on the cluster's feature compatibility version (FCV):

    • For clusters with FCV 9.0 or later, logs return a MongoServerError similar to the following: Replica set writes blocked, reason: InsufficientDiskSpace.

    • For clusters with FCV earlier than 9.0, logs return a MongoServerError similar to the following: User writes blocked, reason: DiskUseThresholdExceeded.

    On Atlas Infinite clusters, Atlas blocks writes when the logical data size on the replica set approaches the cluster's max storage limit. To learn more about storage limits on Atlas Infinite clusters, see Manage Storage on an Atlas Infinite Cluster.

Load Shedding is an IWM protection that rejects incoming database operations as early as possible to limit their impact on the cluster. Unlike IWM's always-on protections, Load Shedding is optional, so you can control whether Atlas applies it to your cluster. When enabled, Load Shedding engages only during sustained overload, after other protections can't relieve enough pressure.

As part of IWM's always-on protections, Atlas runs an ingress request rate limiter on each mongod process in your cluster. This rate limiter admits operations into the process at a controlled rate, and adds any operation requests that arrive faster than this rate to an ingress queue on that process.

When Load Shedding is enabled, Atlas bounds the size of the ingress queue on each mongod process in the cluster. If an operation arrives while a process's queue is full, Atlas rejects the operation with an overload error. This prevents the queue from growing too large during sustained overload.

When Load Shedding is disabled, Atlas doesn't bound the size of the ingress queue on any mongod process in your cluster or reject any operations that arrive faster than the rate limiter's admission rate. During sustained overload, each mongod process continues to add an unbounded number of operations to its queue, so operations take longer to complete.

If reactive auto-scaling is enabled on your cluster, Atlas uses the rate of operations that Load Shedding rejects to determine when to scale your cluster up. Atlas reports this rate as Rejected Operations under Operation Rate Limiting in Atlas metrics. To learn more, see How Atlas Scales Cluster Tier.

Load Shedding is available only on M30+ replica set clusters running MongoDB 9.0 or later. By default, Atlas sets the value of the Load Shedding setting to Atlas Managed, which lets Atlas control the effective state of Load Shedding for your cluster. On MongoDB 9.0, that state is Disabled.

Important

When enabled, Load Shedding can increase the number of overload errors on your cluster. To learn how to handle these errors, see Handle Overload Errors from Load Shedding.

To configure Load Shedding on your cluster, use the Atlas UI or Atlas Administration API:

Intelligent Workload Management (IWM) is active on all Dedicated clusters running MongoDB 9.0 or later. IWM is not available on Free clusters or Flex clusters.

When IWM is active on a cluster, Atlas applies a subset of IWM protections based on the cluster's tier, topology, MongoDB version, and Load Shedding configuration. The following table lists the protections that Atlas applies to each combination of cluster tier and topology on MongoDB 9.0 and later:

IWM protection
M10 and M20 clusters
M30+ replica set clusters
M30+ sharded clusters

Ingress queueing

No

Yes

No

Execution control prioritization

No

Yes

No

Cache eviction and memory tuning

Yes

Yes

Yes

Connection rate limiting

Yes

Yes

Yes

Write blocking

Yes

Yes

Yes

Load Shedding

No

Yes, when enabled

No

Note

Exceptions to MongoDB 9.0 Version Requirement.

The following IWM protections are active on clusters running MongoDB versions earlier than 9.0:

  • Write blocking is active on M10+ clusters running MongoDB 8.0 or later.

  • Connection rate limiting is active on M10 and M20 clusters running the following MongoDB versions:

    • 7.0.23 and later 7.0.x releases

    • 8.0.12 and later 8.0.x releases

    • 8.1.2 and later 8.1.x releases

Atlas provides several observability surfaces that show when IWM protects your cluster and what operations Atlas delayed or rejected: