Autoscaling for Performance: Scaling Policies That Work

Autoscaling is the mechanism that matches compute capacity to workload demand without manual intervention. When configured well, autoscaling maintains consistent performance during traffic spikes while controlling costs during quiet periods. When configured poorly, it either scales too slowly (causing performance degradation) or too aggressively (wasting budget on idle resources).

The gap between theory and practice is large. Default autoscaling configurations rarely work well for production workloads because they optimize for a generic scenario that matches no one's actual traffic patterns. Effective autoscaling requires understanding your application's scaling characteristics, choosing the right metrics to drive scaling decisions, and tuning timing parameters to match your workload's behavior.

Horizontal vs Vertical Scaling

The fundamental choice is whether to scale by adding more instances (horizontal) or by increasing the capacity of existing instances (vertical). Most modern architectures favor horizontal scaling, but the decision depends on application architecture and workload characteristics.

Horizontal Scaling Vertical Scaling 2 CPU 2 CPU 2 CPU +2 CPU (new) Add more instances of same size Total: 6→8 CPU across 3→4 nodes 4 CPU 16GB → 8 CPU 32GB (upgraded) Upgrade instance to larger size Total: 4→8 CPU on 1 node ✓ No downtime ✓ Linear capacity growth ✓ Fault tolerant ✗ Requires stateless design ✗ Load balancer needed ✓ Simple architecture ✓ No state distribution ✗ Downtime during resize ✗ Upper limit on instance size ✗ Single point of failure

When Horizontal Scaling Wins

Horizontal scaling suits stateless services that can distribute load across instances: web servers, API gateways, worker queues, and containerized microservices. The application must handle concurrent instances gracefully, storing session state externally in a cache or database rather than in local memory.

The key advantage is granularity. Horizontal scaling adds capacity in small increments (one container at a time) and removes it just as easily. There is no theoretical upper bound on total capacity, and each instance is independently replaceable, providing natural fault tolerance.

When Vertical Scaling Wins

Vertical scaling suits workloads that benefit from larger single-machine resources: in-memory databases, single-threaded processing tasks, and applications with shared mutable state. A database server that needs more memory for its buffer pool or a batch processing job that needs more CPU cores for parallel computation may be easier to scale vertically than to distribute across multiple instances.

The practical limit is the largest available instance type. Once you hit that ceiling, horizontal scaling or architectural redesign becomes necessary.

Choosing the Right Scaling Metric

The metric that drives scaling decisions determines how quickly and accurately the system responds to demand changes. The wrong metric leads to either premature scaling (wasting resources) or delayed scaling (degrading performance).

CPU Utilization

CPU utilization is the most common scaling metric and works well for compute-bound workloads. Target utilization typically ranges from 60-80%. Below 60% wastes capacity; above 80% leaves insufficient headroom for burst handling and background processes.

# AWS Auto Scaling: CPU-based target tracking aws autoscaling put-scaling-policy \ --auto-scaling-group-name my-asg \ --policy-name cpu-target-tracking \ --policy-type TargetTrackingScaling \ --target-tracking-configuration '{ "PredefinedMetricSpecification": { "PredefinedMetricType": "ASGAverageCPUUtilization" }, "TargetValue": 70.0, "ScaleInCooldown": 300, "ScaleOutCooldown": 60 }'

CPU utilization fails as a scaling metric when the application is I/O-bound rather than compute-bound. A web server waiting on database queries may sit at 15% CPU utilization while serving requests slowly because all threads are blocked on network I/O. Scaling based on CPU in this scenario never triggers, even though more instances could absorb more concurrent connections.

Request Count and Latency

Scaling based on request count or response latency aligns more directly with user experience than infrastructure metrics. If each instance can handle 500 requests per second at acceptable latency, you can scale to maintain that ratio as traffic changes.

Request-based scaling responds to demand changes faster than CPU-based scaling because request count increases immediately when traffic arrives, while CPU utilization lags behind as requests begin processing.

Queue Depth

For worker-based architectures processing jobs from a queue, queue depth is the most natural scaling metric. When the queue grows, more workers are needed; when it shrinks, workers can be released. The target is to keep the queue at or near zero during normal operation, with scaling triggered when the queue depth exceeds a threshold.

Custom Application Metrics

The most effective scaling metrics are often application-specific: active WebSocket connections for a real-time collaboration tool, concurrent video transcoding jobs for a media platform, or pending payment transactions for an e-commerce checkout service. Publishing custom metrics to your cloud provider's monitoring service allows autoscaling to react to the signals that actually predict capacity needs.

Scaling Policy Types

Cloud providers offer several policy types, each with different response characteristics. Understanding the tradeoffs helps you choose the right policy for your workload.

Policy TypeMechanismResponse SpeedBest For
Target TrackingMaintains a metric at target valueModerate (1-3 min)Steady, predictable traffic
Step ScalingAdds/removes fixed amounts at thresholdsFast (alarm-triggered)Multi-level responses
Simple ScalingOne action per alarmSlow (waits for cooldown)Simple, low-risk
Scheduled ScalingPre-set capacity at known timesInstant (pre-provisioned)Known traffic patterns
Predictive ScalingML forecasts future demandProactive (ahead of demand)Recurring daily/weekly patterns

Target Tracking Policies

Target tracking is the recommended default for most workloads. You specify a target value for a metric (such as 70% average CPU utilization), and the autoscaler continuously adjusts capacity to maintain that target. It handles both scale-out and scale-in automatically without requiring separate alarm configurations.

Target tracking works best when the relationship between the scaling metric and capacity is roughly linear. If doubling the instance count halves the CPU utilization, target tracking converges quickly to the right capacity. If the relationship is non-linear (such as when a shared database becomes the bottleneck), target tracking may oscillate.

Step Scaling for Graduated Response

Step scaling allows different actions at different metric thresholds. This is useful for workloads with non-linear scaling requirements or when you want aggressive scale-out but conservative scale-in.

# Step scaling: graduated response to CPU utilization # 70-80% CPU: add 1 instance # 80-90% CPU: add 2 instances # 90%+ CPU: add 4 instances (emergency burst) # # Scale-in: # 50-60% CPU: remove 1 instance # Below 50%: remove 2 instances Step Adjustments: - MetricIntervalLowerBound: 0 # 70% (alarm threshold) MetricIntervalUpperBound: 10 # 80% ScalingAdjustment: 1 - MetricIntervalLowerBound: 10 # 80% MetricIntervalUpperBound: 20 # 90% ScalingAdjustment: 2 - MetricIntervalLowerBound: 20 # 90%+ ScalingAdjustment: 4

Predictive Scaling

Predictive scaling uses machine learning to analyze historical traffic patterns and pre-provision capacity before demand arrives. This eliminates the scale-out delay entirely for predictable workloads with daily, weekly, or seasonal patterns.

Predictive scaling works best alongside reactive policies. The predictive policy handles the expected baseline, while target tracking or step scaling handles unexpected spikes. The two policies do not conflict because the autoscaler always uses the higher of the two capacity recommendations.

Tuning Cool-Down Periods

Cool-down periods prevent rapid oscillation by enforcing a minimum delay between scaling actions. Setting these correctly is critical for both performance and cost.

Scale-Out Cool-Down

The scale-out cool-down determines how long after adding capacity the autoscaler waits before evaluating whether to add more. Set this value based on how long a new instance takes to become fully operational and start absorbing load:

  • Container startup (ECS/Kubernetes): 30-90 seconds typically; set cool-down to 60-120 seconds
  • VM startup (EC2/Compute Engine): 2-5 minutes; set cool-down to 180-300 seconds
  • Application warm-up (JVM, caches): Add warm-up time to the cool-down period

A cool-down that is too short causes over-provisioning: the autoscaler adds instances before the previous batch has started serving traffic, sees utilization still high, and adds more. A cool-down that is too long causes under-provisioning during rapid traffic growth.

Scale-In Cool-Down

The scale-in cool-down should be longer than scale-out (typically 2-5x) to prevent premature capacity reduction. Scale-in is less urgent than scale-out: excess capacity costs money but does not affect user experience, while insufficient capacity directly degrades performance.

A common pattern is asymmetric timing: scale out aggressively (60-second cool-down) and scale in conservatively (300-second cool-down). This bias favors performance over cost, which is the right tradeoff for most user-facing services.

Scaling in Kubernetes

Kubernetes provides its own autoscaling mechanisms that layer on top of (or replace) cloud provider autoscaling. Understanding how these interact is essential for Kubernetes-based deployments.

Horizontal Pod Autoscaler (HPA)

The HPA adjusts the number of pod replicas based on observed metrics. It queries the metrics server every 15 seconds by default and uses a stabilization window to prevent flapping.

apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: api-server spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: api-server minReplicas: 3 maxReplicas: 50 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70 - type: Pods pods: metric: name: requests_per_second target: type: AverageValue averageValue: "500" behavior: scaleUp: stabilizationWindowSeconds: 30 policies: - type: Percent value: 50 periodSeconds: 60 scaleDown: stabilizationWindowSeconds: 300 policies: - type: Pods value: 2 periodSeconds: 60

Cluster Autoscaler

The HPA scales pods, but pods need nodes to run on. The Cluster Autoscaler adds and removes nodes based on pending pod requests. When a pod cannot be scheduled due to insufficient node resources, the Cluster Autoscaler provisions a new node. When a node is underutilized and its pods can be moved elsewhere, the node is drained and terminated.

The interaction between HPA and Cluster Autoscaler creates a two-stage scaling pipeline: HPA requests more pods, some pods become pending due to insufficient node capacity, the Cluster Autoscaler adds nodes, pods are scheduled onto the new nodes. This pipeline adds latency: typical end-to-end scale-out takes 2-5 minutes in Kubernetes compared to 30-60 seconds for pod-only scaling on existing nodes.

Cost Optimization Strategies

Autoscaling directly impacts cloud costs. The goal is to run at the minimum capacity that meets performance targets, with enough headroom to handle variability without triggering scaling events for every fluctuation.

Right-Sizing Base Capacity

Analyze your traffic patterns to determine the minimum capacity you need at all times. Set your autoscaling group's minimum size to this baseline. Ensure the baseline can handle typical non-peak traffic without triggering scale-out, because every unnecessary scaling event adds latency as new instances warm up.

Spot and Preemptible Instances

Use spot instances for the scale-out portion of your capacity. Your base capacity runs on reliable on-demand or reserved instances, while burst capacity uses spot instances at 60-90% cost savings. The risk of spot interruption is acceptable for burst capacity because losing those instances only reduces you to your baseline, which is sized to handle normal traffic.

Scheduled Scaling for Known Patterns

If your traffic has predictable peaks (business hours, weekly cycles, marketing campaign launches), use scheduled scaling to pre-provision capacity before the peak arrives. This avoids the performance cost of reactive scaling during the critical ramp-up period and ensures the first users in the peak experience full performance.

Common Autoscaling Pitfalls

Scaling on the Wrong Metric

Using CPU utilization to scale an I/O-bound application is the most common mistake. If your application spends most of its time waiting on database queries, network calls, or disk I/O, CPU-based scaling never triggers even though the application is at capacity. Use request rate, concurrent connections, or application-specific metrics instead.

Insufficient Minimum Capacity

Setting the minimum to 1 instance means a single instance must handle traffic while scale-out is triggered and new instances are provisioning. For any production workload, maintain at least 2-3 instances as the minimum for redundancy and to distribute load during the scale-out delay.

Missing Health Checks

New instances must pass health checks before receiving traffic. Without proper health checks, the load balancer sends requests to instances that are still starting up, causing errors and degraded Core Web Vitals. Configure both startup probes (allowing time for initialization) and readiness probes (confirming the application can serve requests) before marking an instance as healthy.

Ignoring Scale-In Behavior

Uncontrolled scale-in can terminate instances that are actively processing requests. Enable connection draining to finish in-flight requests before terminating an instance. In Kubernetes, use preStop hooks and terminationGracePeriodSeconds to give pods time to complete work.

Frequently Asked Questions

What is the ideal CPU target for autoscaling?
For most web workloads, target 60-70% average CPU utilization. This leaves enough headroom to absorb short bursts without triggering scale-out while keeping instances productive. Compute-intensive workloads like video encoding can target higher (75-85%) because their CPU usage is more predictable and less bursty.
How fast should autoscaling respond to traffic spikes?
The target is to have new capacity ready before users experience degradation. For container-based deployments, aim for under 2 minutes from spike detection to new instances serving traffic. For VM-based deployments, 3-5 minutes is typical. If this is too slow for your workload, combine reactive scaling with predictive scaling or maintain higher baseline capacity.
Should I use horizontal or vertical autoscaling?
Use horizontal scaling for stateless applications like web servers, API services, and worker processes. Use vertical scaling for stateful workloads where distributing data across instances is impractical, such as in-memory databases or legacy monoliths. Many cloud providers now support automatic vertical scaling for Kubernetes pods through the Vertical Pod Autoscaler.
How do I prevent autoscaling from oscillating?
Oscillation occurs when scale-out reduces the metric below the scale-in threshold, causing capacity removal, which pushes the metric back above the scale-out threshold. Prevent this by using stabilization windows (waiting several minutes before acting on metric changes), setting asymmetric cool-downs (longer for scale-in), and ensuring scale-in thresholds are well below scale-out thresholds.
Can I combine multiple scaling policies?
Yes, most cloud platforms support multiple concurrent scaling policies. When multiple policies recommend different capacities, the highest value wins to ensure performance is not compromised. A common combination is predictive scaling for baseline traffic plus target tracking for unexpected spikes, with scheduled scaling for known events like product launches.