17 November 2025 · 5 min read

Autoscaling buys you the next spike

Reactive autoscaling cannot serve the spike that triggers it. The new capacity arrives after a chain of delays that add up to minutes. What it buys is the spike after this one.

The promise of autoscaling is that capacity appears when demand does. The mechanism cannot keep that promise, because capacity is a consequence of demand being measured, and measurement takes time at every step. By the time a reactive policy has noticed a spike, decided about it, launched an instance, waited for it to boot, and let it into the load balancer, the spike that triggered all of that is several minutes old. What the new instance serves is whatever comes next. That is not a flaw to be tuned away. It is what "reactive" means, and it changes what the policy is for.

The arrival gap

I call the total delay the arrival gap: the time between demand crossing the threshold and a new instance taking its first request. It is a sum of documented settings and physical facts, and every one of the settings defaults to a number that assumes you are not in a hurry.

First the metric has to be observed. CloudWatch collects EC2 metrics in periods of a minute with detailed monitoring and five minutes without it, so the earliest the system can know about a change is the end of the current period. Then the alarm has to evaluate: a target tracking policy, per the AWS documentation, creates alarms that need several consecutive periods above the target before they fire, so that a single noisy minute does not launch a fleet. Then an instance has to launch, which takes as long as the image takes to boot and the bootstrap takes to run. Then the load balancer's health check has to pass, after its own grace period. And then, with a default instance warmup configured, the instance is held out of the aggregate metrics for that long; AWS's warm-up and cooldown page explains how the warmup falls back to the default cooldown, which is 300 seconds, when it is not set.

The arrival gap as a sum of stages One stacked horizontal bar in seconds: metric period 60, alarm evaluation 180, instance launch and bootstrap 120, health check grace 60, instance warmup 300, for a total of about 720 seconds or twelve minutes. Warmup is the largest segment. Twelve minutes, from documented defaults and a typical boot Seconds from demand crossing the threshold to the new instance serving alarm 180 boot 120 warmup 300 60 60 Metric 60, alarm 180, launch and bootstrap 120, health grace 60, warmup 300. Only the boot figure is a guess; the rest are defaults or usual values. total: about 720 seconds
Source: the 300-second default cooldown and warmup fallback from the EC2 Auto Scaling warm-up and cooldown documentation; the metric period and evaluation from the target tracking documentation; boot and grace are illustrative values stated as assumptions.

Twelve minutes is a long time. Most of the spikes a small service sees are shorter than that: a link shared somewhere, a batch job that fires on the hour, a scheduled email that sends everyone to the site at once. Against a spike that lasts five minutes, a twelve-minute arrival gap means the new capacity comes online after the spike is over, serves the trough, and is then scaled back in by the same policy. The requests that failed during the spike failed. The policy was decoration.

What it does buy

Reactive scaling is not useless. It is late, which is different, and late is fine for the demand it was designed for: a rise that lasts long enough for the arrival gap to be a small fraction of it. A morning ramp that builds over an hour, a campaign that doubles traffic for a week, a slow growth that the fleet has to follow across a quarter. Against those, the gap is a few minutes of under-provisioning at the front of a long period of correct provisioning, and the policy earns its keep.

Demand against reactive capacity on one timeline A demand curve rises sharply at one point and falls back a few minutes later. A capacity line stays flat, then steps up twelve minutes after the rise, after the demand has already fallen. The window between the rise and the capacity step is shaded as the period in which requests fail. The step arrives after the spike A short spike under a reactive policy; shaded where requests fail demand capacity 0 spike +12 min +30 min demand above capacity capacity above demand
Illustrative: the timing relationship the arrival gap produces, drawn to show shape rather than measured traffic.

For everything shorter, the honest question is what the policy is defending against, and the answer is the next spike. If spikes come in waves, one after another, the capacity that arrived late for the first one is in place for the second, and that is worth something. If spikes come alone, it is worth nothing, and the money spent on the extra instance that arrives late and leaves later is money that could have bought a spare instance that was already there.

Scaling in is where the money goes

The same delays run in the other direction, and they are set longer on purpose. Target tracking scales in more cautiously than it scales out, with more evaluation periods before it removes an instance, because removing capacity too early is the failure operators notice. The consequence is that a fleet that scaled out late for a spike also scales in late after it, and the extra instance that arrived when the spike was over stays for a good while after that. Over a day of isolated spikes, the fleet spends most of its time one instance larger than the traffic needs and never one instance larger at the moment it needed to be. That is the cost profile of reactive scaling against short spikes: paying for capacity in the troughs and lacking it at the peaks.

Seen that way, the spare instance is not an extravagance. It is the same instance the policy would have bought anyway, held at a time when it is useful instead of at a time when it is not.

When the schedule beats the policy

The QR ticketing system I built serves visitor attractions, and an attraction's demand is not a mystery. The gates open at a fixed time. The scan-in traffic starts at that minute, peaks in the first hour, and follows the day's programme. Nobody needs a metric to discover that ten o'clock is busier than nine; it is printed on the sign outside. A reactive policy would spend the first twelve minutes of every morning learning something the operator has known for years.

The same is true of most of the small services I have run. A CRM's traffic follows office hours in the client's time zone. A marketing site's traffic follows the campaign calendar. A helpdesk assistant's traffic follows the shifts of the people who ask it questions. In each case someone in the business can draw the demand curve from memory, and a policy that has to rediscover it from a metric every day is doing expensive, slow, slightly wrong work that a calendar entry does for free.

For a service whose peaks are on a schedule, two things beat target tracking, and both are cheaper. One spare instance, N plus one, absorbs any spike smaller than one instance's worth of capacity with no gap at all, because it is already there. A scheduled action, which AWS supports directly, raises the fleet before the known peak and lowers it after, with an arrival gap of zero because the launch happens before the demand rather than after it. Reactive scaling then sits behind both as a backstop for the day the sign outside is wrong.

An opening-hour surge against three capacity strategies Demand steps up sharply at opening time and stays high for an hour. Reactive capacity steps up twelve minutes late. N plus one capacity is flat and slightly above demand from the start. Scheduled capacity steps up fifteen minutes before opening. Only the reactive line has a window where demand exceeds capacity. Three ways to meet ten o'clock An opening-hour step in demand; a model with the construction stated below 9:30 10:00 10:30 11:00 Time of day demand reactive scheduled N plus one reactive gap
Illustrative: demand modelled as a step at opening, reactive capacity as the same step delayed by the twelve-minute arrival gap, scheduled capacity as the step fifteen minutes early, and N plus one as a flat line above the peak.

Computing your own gap

The arrival gap is a number you can compute from your own settings in a few minutes, and it is worth computing before choosing a policy rather than after an incident. Read the metric period from the monitoring level. Read the alarm's evaluation periods from the policy. Time a launch from the console to the first successful health check. Add the warmup or the cooldown it falls back to. The sum is how late every reactive decision will be, and it should be compared with the length of the spikes in your own request logs, not with a feeling about how fast the cloud is.

If the gap is shorter than your typical spike, reactive scaling will catch most of them, and the tuning that matters is bringing the gap down further: detailed monitoring, fewer evaluation periods, a faster image, a shorter warmup. If the gap is longer than your typical spike, no tuning of the policy fixes the problem, because the policy is the wrong tool. The tools for that case are a spare instance and a calendar, and for a great many small services they are the whole answer. Autoscaling then does what it is good at, which is following the slow tide, and stops being asked to catch the wave.

AWSAutoscalingCapacity
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS