This page covers cluster tier auto-scaling on Atlas Infinite clusters, which support both reactive and predictive cluster tier auto-scaling. To learn how each mechanism works, see Reactive Auto-Scaling for Cluster Tier and Predictive Auto-Scaling for Cluster Tier.
Cluster Tier Auto-Scaling
Atlas uses reactive and predictive auto-scaling for cluster tiers. Atlas chooses its auto-scaling mechanism based on your cluster's type, tier, and workload pattern.
Reactive auto-scaling. Atlas uses thresholds, and not prediction, to trigger scaling events based on current resource usage. Reactive auto-scaling occurs after sustained high or low resource usage. To learn more, see Reactive Auto-Scaling for Cluster Tier.
Predictive auto-scaling. Atlas uses machine learning to anticipate future scaling needs based on historical usage patterns and attempts to trigger scaling events before the forecasted workload spike arrives.
Predictive auto-scaling is an extension of cluster tier auto-scaling and falls back to reactive auto-scaling. Atlas continues to rely on reactive auto-scaling to manage unexpected spikes in workload that aren't cyclical or predictable. Atlas uses predictive auto-scaling for eligible clusters. To learn more, see Predictive Auto-Scaling for Cluster Tier.
Reactive Auto-Scaling for Cluster Tier
Note
Auto-Scaling Term Usage
In all of Atlas documentation, whenever the auto-scaling term is used without the word "predictive", it refers to the reactive auto-scaling mechanism. See also predictive auto-scaling.
You can configure the cluster tier ranges that Atlas uses to automatically scale your cluster tier in response to cluster usage.
To optimize resource utilization and improve cost profile, Atlas reactive auto-scaling detects sustained higher demand and short-term peak traffic and adjusts cluster tier based on real-time resource usage.
To help control costs, you can specify a range of maximum and minimum cluster sizes that your cluster can automatically scale to.
Reactive auto-scaling works on a rolling basis, and the process doesn't incur any downtime. Atlas maintains a primary node during this process, but the nodes are upgraded one-by-one and are unavailable while being upgraded.
To learn about recommendations for scalability, including avoiding resource drift when using infrastructure as code tools with reactive auto-scaling, see Recommendations for Atlas Scalability in the Atlas Architecture Center.
Eligible Clusters for Reactive Auto-Scaling
Atlas cluster tier reactive auto-scaling is available for all dedicated Atlas Core cluster tiers under the General and the Low-CPU cluster classes. Reactive auto-scaling is also available for Atlas Infinite clusters.
How Atlas Scales Cluster Tier
On Atlas Infinite clusters, scaling the cluster tier also changes the storage IOPS and storage throughput available to the cluster. To review the storage performance values for each tier, see Atlas Infinite Storage IOPs and Throughput Values per Cluster Tier.
Note
You can't update the storage performance values of a given tier. To change storage performance, change the cluster tier.
The local storage on a compute node is separate from cluster storage. To learn how much local storage each cluster tier provides for query spilling, see .
Atlas relies on host ping data for autoscaling decisions. Dedicated cluster data nodes continuously send this ping data to the control plane regardless of whether autoscaling is enabled. When you enable autoscaling, Atlas can use this historical data to scale immediately if scaling conditions are met.
Atlas scales your cluster to another tier in the same class. For example, Atlas scales General clusters to other General cluster classes, but doesn't scale General clusters to Low-CPU cluster classes.
Atlas won't scale your cluster tier if the new cluster tier would fall outside of your specified Minimum and Maximum Cluster Size range.
If you deploy read-only nodes and want your cluster to scale faster, consider adjusting your Replica Set Scaling Mode.
The exact reactive auto-scaling criteria are subject to change in order to ensure appropriate cluster resource utilization.
Important
For dedicated Atlas Core clusters, if you restore a snapshot with a larger size than the storage capacity of the destination cluster, the cluster does not automatically scale.
Atlas uses the following resource utilization and operation admission control concepts to determine when to scale your cluster up or down:
Absolute System CPU Utilization: Total CPU usage of all processes on the node. Visible as System CPU in Atlas metrics.
Relative System CPU Utilization: Value Atlas uses for autoscaling decisions on
M10andM20clusters. This is calculated as:Relative System CPU Utilization = Normalized System CPU / Baseline CPU Utilization Where:
Normalized System CPU: Total CPU usage summed across all cores, normalized to the baseline CPU utilization. Visible as Normalized System CPU in Atlas metrics.
Baseline CPU Utilization: Fraction of full CPU guaranteed to your instance by the cloud provider, typically 20%-50% for burstable instance types. Not visible in Atlas metrics. To learn more, see baseline CPU utilization.
For example, using 20% as a lower-end estimate for the Baseline CPU Utilization, the following Relative System CPU Utilization values correspond to these Normalized System CPU values in Atlas metrics:
75%relative system CPU utilization equals15%Normalized System CPU (75% of 20%).90%relative system CPU utilization equals18%Normalized System CPU (90% of 20%).
Atlas caps the Relative System CPU Utilization at 100%, even when the calculation exceeds it. If a scale-up occurs when Normalized System CPU appears low, contact MongoDB Support.
System Memory Utilization: Total memory usage across all processes on the node, expressed as a percentage of total memory available to the node. This is calculated as:
System Memory Utilization = Memory Used / Total Memory * 100 Where:
Memory Used: Number of bytes of physical memory currently in use on the host. Visible as System Memory: Memory Used (bytes) in Atlas metrics.
Total Memory: Total physical memory available to the node, as reported by the operating system. Atlas does not display this value as a separate metric. Not visible in Atlas metrics.
Note
The System Memory Utilization value that Atlas uses for autoscaling decisions might differ slightly from the value shown in the Atlas metrics panel. If a scale-down does not occur when System Memory Utilization appears low, contact MongoDB Support.
Queued or Rejected Operations: Combined rate of operations that Atlas queues or rejects to protect your cluster from overload as part of Intelligent Workload Management (IWM). Atlas calculates this as:
Queued or Rejected Operations = Queued Operations + Rejected Operations Where:
Queued Operations: Average rate per minute of incoming operations that Atlas adds to the ingress request rate limiter queue to wait for admission into the cluster. The per-second rate is visible as Operation Rate Limiting: Queued operations in Atlas metrics.
Rejected Operations: Average rate per minute of incoming operations that Atlas rejects because the cluster is overloaded and Load Shedding is active. The per-second rate is visible as Operation Rate Limiting: Rejected operations in Atlas metrics.
To learn how IWM works to queue or reject operations in response to cluster overload, see Intelligent Workload Management.
A combined value above zero means that your cluster is under overload.
The following sections describe how Atlas uses these metrics to determine when to scale your cluster up or down.
Conditions for Scaling Up
To manage dynamic workloads for your applications, Atlas reactively scales up nodes in your cluster under the conditions described in this section.
To achieve optimal resource utilization and cost profile, Atlas avoids scaling up the cluster to the next tier if:
The
M10orM20cluster has been scaled up in the past 20 minutes or one hour, depending on thresholds.The
M30+cluster has been scaled up in the past 10 minutes or one hour, depending on thresholds.The cluster has been scaled up in the past 10 minutes, for the Queued or Rejected Operations criterion.
For example, if the cluster tier has not been changed since 12:00, Atlas will scale an M30+ cluster at 12:10, if the cluster's current normalized System CPU Utilization is greater than 90%.
If the next cluster tier is within your Maximum Cluster Size range, Atlas scales operational nodes in your cluster up to the next tier if at least one of the following criteria is true for any cluster node of this type.
Note
The conditions in this section describe operational nodes. For analytics nodes on any cloud provider, Atlas scales them up to the next tier if the average Normalized System CPU or the System Memory Utilization has exceeded 75% of resources available to any cluster node for the past one hour. Atlas doesn't apply the Queued or Rejected Operations criterion to analytics nodes.
The following list groups the criteria by cluster tier. Within each tier, CPU-related criteria appear first, followed by memory-related criteria. Within each of those two sets, criteria specific to a cloud provider appear first. The remaining criteria appear in order from most restrictive to least restrictive. The overload criterion applies to every dedicated tier and appears last.
M10andM20clusters:AWS. The average normalized Relative System CPU Utilization has exceeded 90% for the past 20 minutes and the average non-normalized Absolute System CPU Utilization for CPU steal has exceeded 30% for the past 3 minutes.
Azure. The average normalized Relative System CPU Utilization has exceeded 90% for the past 20 minutes and the average non-normalized Absolute System CPU Utilization for softIRQ has exceeded 10% for the past 3 minutes.
The average normalized Absolute System CPU Utilization has exceeded 90% of resources available to the cluster for the past 20 minutes.
The average normalized Relative System CPU Utilization has exceeded 75% of resources available to the cluster for the past one hour.
The average System Memory Utilization has exceeded 90% of resources available to the cluster for the past 10 minutes.
The average System Memory Utilization has exceeded 75% of resources available to the cluster for the past one hour.
Note
If a scale-up occurs when Normalized System CPU appears low, contact MongoDB Support.
M30+clusters:The average Normalized System CPU has exceeded 90% of resources available to the cluster for the past 10 minutes.
The average Normalized System CPU has exceeded 75% of resources available to the cluster for the past one hour.
The average System Memory Utilization has exceeded 90% of resources available to the cluster for the past 10 minutes.
The average System Memory Utilization has exceeded 75% of resources available to the cluster for the past one hour.
All Dedicated clusters,
M10+:Queued or Rejected Operations remains above zero for every sample in the past 10 minutes.
Atlas requires the rate to stay above zero for the entire 10-minute window rather than averaging it, so a brief burst of queueing or rejecting operations during a short traffic spike doesn't scale your cluster. Sustained load shedding indicates that your workload exceeds what the current tier can admit, so Atlas scales up to relieve the overload.
Note
This criterion measures whether Atlas shed load, not how much. Any sustained rate above zero meets the threshold.
These thresholds ensure that your cluster scales up quickly in response to high loads, maintaining its performance and reliability.
Note
Atlas does not trigger cluster tier auto-scaling during a simulated regional outage. This behavior may also occur during an actual regional outage if the cluster has insufficient healthy nodes to support scaling operations.
Important
Sudden Workload Spikes
Scaling up to a greater cluster tier requires enough time to prepare backing resources. Automatic scaling may not occur when a cluster receives a burst of activity, such as a bulk insert. To reduce the risk of running out of resources, plan to scale up clusters before bulk inserts and other workload spikes.
Example
Consider an example scenario with the following values to see how Atlas evaluates the scaling conditions. The baseline CPU utilization is not visible in the Atlas metrics panel and can range from 20%-50% for burstable instance types. You can use any value in that range to estimate upper and lower bounds, and this example uses 20% as the lower end of that range.
Normalized System CPU: 60%
CPU steal: 10%
Evaluating the conditions:
Condition 1 (AWS): Requires average Relative System CPU Utilization > 90% for 20 minutes AND average CPU steal > 30% for 3 minutes.
Relative CPU: 60% ÷ 20% = 300%, capped at 100%. First threshold met.
CPU steal is 10%, which doesn't exceed 30%. Second threshold not met.
Result: Condition 1 Not Met. Both thresholds must be true.
Condition 2: Requires average Normalized System CPU > 90% for 20 minutes.
Normalized System CPU is 60%, which doesn't exceed 90%.
Result: Condition 2 Not Met.
Condition 3: Requires average Relative System CPU Utilization > 75% for 1 hour.
Relative CPU: 60% ÷ 20% = 300%, capped at 100%.
Result: Condition 3 Met. Atlas triggers auto-scaling.
Conditions for Scaling Down
To optimize costs, Atlas reactively scales down nodes in your cluster under the conditions described in this section.
Atlas begins checking these conditions from the moment you enable downscaling, not retroactively. Even if your cluster met these conditions before you enabled downscaling, Atlas will not scale down until the required time windows have elapsed since you enabled the feature.
If the next lowest cluster tier is within your Minimum Cluster Size range, Atlas scales the nodes in your cluster down to the next lowest tier if all of the following criteria are true for all nodes in the cluster:
All nodes:
Atlas hasn't scaled the cluster down (manually or automatically) in the past 24 hours.
Atlas hasn't provisioned or unpaused the cluster in the past 24 hours.
Atlas hasn't stopped and restarted any cluster nodes in the past 12 hours.
The average Normalized System CPU is below 45% of resources available to the cluster over at least the last 10 minutes AND the last 4 hours. Atlas uses the "4 hours average" checkpoint as an indication that the CPU load has settled down on the observed level. Atlas uses the "10 minutes average" checkpoint as an indication that no recent CPU spikes have occurred that Atlas didn't capture with the "4 hour average" checkpoint.
Note
For
M10andM20tiers, Atlas applies the 45% CPU threshold relative to the instance baseline CPU utilization rather than the standard 100% baseline. Using 20% as a lower-end estimate, the effective absolute CPU threshold for scale-down is approximately 9% (45% of 20%).The average WiredTiger cache usage is below 90% of the maximum WiredTiger cache size for at least the last 10 minutes AND the last 4 hours at the current cluster tier size. This indicates to Atlas that the current cluster isn't overloaded.
The Projected Memory Utilization at the new lower cluster tier is below 60% for at least the last 10 minutes AND the last 4 hours.
To calculate Projected Memory Utilization, Atlas starts with the current memory usage, visible as System Memory: Memory Used (bytes) in Atlas metrics. Atlas subtracts the current WiredTiger cache usage, adds 80% of the maximum WiredTiger cache size on the new lower tier, then divides the result by that tier's total RAM.
This value differs from System Memory Utilization, which measures all memory in use against the RAM on the current tier.
Note
Atlas includes the WiredTiger cache in this calculation to make it more likely that clusters with a full cache, but otherwise low traffic, scale down. Scaling down requires both of the following thresholds to pass:
90%: The current tier's WiredTiger cache usage must be below 90% of its maximum size.
60%: The Projected Memory Utilization on the new lower tier must be below 60%.
These conditions ensure that Atlas scales down operational nodes in your cluster to prevent high utilization states.
Note
Atlas evaluates memory-based scaling using projected memory utilization, which differs from the System Memory Utilization shown in the Atlas UI. If a scale-down does not occur when System Memory Utilization appears low, contact MongoDB Support.
- The average Normalized System CPU and System Memory Utilization over the past 24 hours is below 50% of resources available to the cluster.
Note
M10andM20clusters use lower thresholds to account for caps on CPU usage set by cloud providers after burst periods. These thresholds vary depending on your cloud provider and cluster tier.
Predictive Auto-Scaling for Cluster Tier
Predictive auto-scaling is an extension of auto-scaling.
Atlas uses demand-forecasting for host resource utilization and performs preemptive scaling up of your cluster compute to ensure optimal resource utilization. With predictive auto-scaling, Atlas attempts to scale up your cluster proactively, ahead of cyclical workload spikes.
Predictive auto-scaling is powered by a machine learning model based on historical patterns. Atlas analyzes resource utilization on the primary node to make scaling decisions. The model predicts when resource utilization will be high based on historical usage patterns, and Atlas scales the cluster up if the model forecasts high resource utilization. MongoDB updates the model and its criteria continuously to optimize Atlas performance.
The model analyzes a rolling 4-week input window to identify cyclical patterns. Any pattern observable within this window, for example, hourly, daily, weekly, or bi-weekly cycles, can be captured. Patterns with longer periods, such as monthly or quarterly cycles, fall outside the 4-week window and are not detectable.
Note
For patterns near the upper limit of the window, accuracy may decrease because fewer complete cycles occur within the window.
Predictive auto-scaling has the following benefits for clusters with predictive, cyclical workloads:
Automatically scale up your cluster for cyclical workload patterns within the 4-week input window.
Maintain consistent performance and availability during predictable high-demand periods.
Reduce manual scaling tasks or scheduled scripts by letting Atlas manage capacity increases.
Seamlessly fall back to reactive auto-scaling when changes to the cluster workload fall outside of predictable patterns and are non-cyclical or unpredictable.
To trigger predictive auto-scaling, your cluster must maintain continuous activity logs for two weeks. Once it meets this criterion, the system enables predictive auto-scaling.
Note
If you pause your cluster, predictive auto-scaling requires two consecutive weeks of activity before it can resume.
Behavior of Predictive Auto-Scaling
The following statements describe how predictive auto-scaling works:
Atlas attempts to scale up your cluster instance size before the forecasted load arrives.
When Atlas scales your cluster predictively based on forecasted metrics, it can scale up by at most two tiers at a time.
Predictive auto-scaling applies only to compute, not storage.
Predictive auto-scaling respects existing auto-scaling minimum and maximum instance sizes.
In cases when Atlas can't use predictive auto-scaling to scale up the cluster, it falls back to using reactive auto-scaling.
Predictive auto-scaling only supports upscaling. There is no predictive down-scaling. Atlas uses reactive auto-scaling to automatically scale down the cluster when the workload decreases.
If predictive up-scaling is scheduled to happen within the next 1 hour, Atlas skips reactive down-scaling.
Eligible Clusters for Predictive Auto-Scaling
Atlas uses predictive auto-scaling for eligible clusters. Eligible clusters for predictive auto-scaling must meet all of the following criteria:
Belong to General and Low-CPU cluster classes.
Have a tier that is
M30or greater.Have auto-scaling enabled. If you enable down-scaling, auto-scaling minimum instance size must be equal to or greater than
M30.Have been active for at least two weeks.
Atlas Infinite clusters at the M30 tier and larger run on AWS Gen2 hardware. Predictive auto-scaling supports these clusters.
In addition, the following criteria affect whether Atlas uses predictive auto-scaling for an eligible cluster:
Predictive auto-scaling applies only to electable and read-only nodes. Atlas doesn't use predictive auto-scaling for search or analytics nodes.
Predictive auto-scaling might not be able to predict non-cyclical and highly dynamic workload spikes in any eligible cluster. In these cases, Atlas relies on reactive auto-scaling.
Considerations for Downward Scaling of Cluster Tier
You can manually reduce your cluster tier from the Edit Cluster page. The following considerations apply when you manually scale down your cluster tier:
Estimate your deployment's range of workloads and then set the Minimum Cluster Size value to the cluster tier that has enough capacity to handle your deployment's workload. Account for any possible spikes or dips in cluster activity.
You can't scale to a cluster tier smaller than
M10.
Configure Auto-Scaling Options
You can configure auto-scaling options when you create or modify a cluster. Atlas recommends compute auto-scaling for Atlas Infinite clusters. You choose whether to use it when you create or modify a cluster.
You can do one of the following:
Review and adjust the upper and lower cluster tiers that Atlas should use when auto-scaling your cluster, or
Opt out of using auto-scaling.
Atlas displays auto-scaling options in the Auto-scale section of the cluster builder for General and Low-CPU tier clusters.
Enable Auto-Scaling with Atlas CLI and Atlas Administration API
You can enable compute auto-scaling when you create or update a cluster using the Atlas CLI or Atlas Administration API. The following examples show how to enable compute auto-scaling for both electable nodes and analytics nodes. Replace the cluster tiers and provider settings with the ones you need.
To configure auto-scaling with the Atlas CLI, create a JSON file that contains the auto-scaling configuration and then specify it in the atlas api clusters updateCluster command.
Use the atlas api clusters updateCluster command to directly call the API and enable auto-scaling settings on an existing cluster. To enable auto-scaling when creating a new cluster, use the atlas api clusters createCluster command.
Create the payload file.
Create a payload.json file with the following content. Replace the placeholder values with your specific cluster configuration:
{ "replicationSpecs": [ { "regionConfigs": [ { "providerName": "{CLOUD-PROVIDER}", "regionName": "{REGION-NAME}", "priority": 7, "electableSpecs": { "instanceSize": "{INSTANCE-SIZE}", "nodeCount": 2 }, "analyticsSpecs": { "instanceSize": "{ANALYTICS-INSTANCE-SIZE}", "nodeCount": 1 }, "autoScaling": { "compute": { "enabled": true, "scaleDownEnabled": true, "minInstanceSize": "{MIN-INSTANCE-SIZE}", "maxInstanceSize": "{MAX-INSTANCE-SIZE}" } }, "analyticsAutoScaling": { "compute": { "enabled": true, "scaleDownEnabled": true, "minInstanceSize": "{MIN-ANALYTICS-INSTANCE-SIZE}", "maxInstanceSize": "{MAX-ANALYTICS-INSTANCE-SIZE}" } } } ] } ] }
Run the update command.
After creating the payload.json file, run the following command to enable auto-scaling on an existing cluster and specify the JSON file with the --file flag:
atlas api clusters updateCluster \ --version 2024-10-23 \ --clusterName {CLUSTER-NAME} \ --groupId {GROUP-ID} \ --file payload.json
You can use the Atlas Administration API to enable auto-scaling by specifying the auto-scaling configuration in the request body.
Use the Update One Cluster in One Project endpoint to enable auto-scaling by including the autoScaling object.
Note
This curl command uses a service account access token (OAuth 2.0) to authenticate instead of API keys. To learn more, see Get Started with the Atlas Administration API.
curl --header "Authorization: Bearer {ACCESS-TOKEN}" \ --header "Accept: application/vnd.atlas.2025-03-12+json" \ --header "Content-Type: application/json" \ --include \ --request PATCH "https://cloud.mongodb.com/api/atlas/v2/groups/{GROUP-ID}/clusters/{CLUSTER-NAME}" \ --data '{ "replicationSpecs": [ { "regionConfigs": [ { "providerName": "{CLOUD-PROVIDER}", "regionName": "{REGION-NAME}", "priority": 7, "electableSpecs": { "instanceSize": "{INSTANCE-SIZE}", "nodeCount": 2 }, "autoScaling": { "compute": { "enabled": true, "scaleDownEnabled": true, "minInstanceSize": "{MIN-INSTANCE-SIZE}", "maxInstanceSize": "{MAX-INSTANCE-SIZE}" } }, "analyticsSpecs": { "instanceSize": "{ANALYTICS-INSTANCE-SIZE}", "nodeCount": 1 }, "analyticsAutoScaling": { "compute": { "enabled": true, "scaleDownEnabled": true, "minInstanceSize": "{MIN-ANALYTICS-INSTANCE-SIZE}", "maxInstanceSize": "{MAX-ANALYTICS-INSTANCE-SIZE}" } } } ] } ]'
Review the Cluster Tier Auto-Scaling Options
To review the enabled auto-scaling options for cluster tier:
Opt Out of Cluster Tier Auto-Scaling
To opt out of cluster auto-scaling (increasing the cluster tier), when creating a new cluster, navigate to the Cluster Tier menu, and un-check the Compute Auto-Scale checkbox in the Auto-scale section.
To opt out of cluster auto-scaling (decreasing the cluster tier), when creating a new cluster, navigate to Cluster Tier menu, and un-check the Allow cluster to be scaled down checkbox in the Auto-scale section.