31 May 2026 · 5 min read

Drift is a speed problem

Drift is treated as a discipline failure: someone touched the console. It is a race. People use the console when code is slower than the emergency, so the fix is in the pipeline.

Every infrastructure-as-code team has the same conversation after a drift incident. Someone changed a security group in the console during an outage, the Terraform state no longer matched the world, and the next apply either reverted the emergency fix or refused to run. The conclusion is always the same: lock the console, make it read-only, route everything through code. It treats drift as a discipline failure. I think it is better modelled as a race, and the race has a measurable winner. People edit the console when the code path is slower than the emergency. Drift accumulates at a rate set by how long a change takes to get from a pull request to production, and locking the console removes the symptom while leaving the cause exactly where it was.

The numbers behind the conversation

The industry surveys describe the symptom well. Firefly's State of IaC 2025 found that 89 percent of respondents had adopted infrastructure as code and only 6 percent had complete coverage, with less than a third continuously monitoring for drift and the rest addressing it reactively. The 2026 edition found a third of respondents had tied drift to a costly production incident, a further 8 percent to significant downtime, and nearly a fifth with no detection or remediation process at all, while 90 percent said their orchestration fell short of what they needed. The gap between adoption and coverage is where drift lives, and the orchestration complaint is the clue to why.

Infrastructure as code adoption, coverage and drift, from two industry surveys Horizontal bars in percent of respondents: adopted infrastructure as code, 89; say their orchestration falls short, 90; tied drift to a costly production incident, 33; continuously monitor drift, under 33; have no drift detection or remediation, nearly 20; have complete coverage of their cloud in code, 6. The complete-coverage bar is highlighted. Adopted almost everywhere, complete almost nowhere Percent of surveyed infrastructure professionals, Firefly State of IaC 2025 and 2026 Adopted IaC 89 Orchestration falls short 90 Drift caused a costly incident 33 Continuously monitor drift under 33 No drift detection at all nearly 20 Complete coverage in code 6
Source: Firefly's State of IaC 2025 (adoption, coverage, monitoring) and State of IaC 2026 (incidents, detection, orchestration), as published on the report pages.

The race

Consider a change that has to happen now: a firewall rule during an attack, a capacity bump during a traffic spike, a certificate that expired at four in the morning. There are two routes. One is the console: open it, click, done, thirty seconds. The other is the code path: edit the module, open a pull request, wait for a plan, get a review, merge, wait for the apply to reach production. On a good day that is twenty minutes. On a bad day, with a queued pipeline, a locked state file, a reviewer in another timezone and a plan that touches something unrelated, it is hours.

Two routes for the same change, with typical latencies and the point where drift enters Top route, the console: open, click, done, about thirty seconds; drift enters here because the state no longer matches. Bottom route, the code path: edit, pull request, plan, review, merge, apply, twenty minutes on a good day and hours on a bad one. A note says the console wins whenever the emergency is shorter than the code path. Whichever route is faster than the emergency wins CONSOLE open click done: 30 s drift enters here CODE PATH edit PR, plan review merge apply done 20 minutes on a good day; hours with a queued pipeline or an absent reviewer. the gold boxes are where the latency lives, and where the fix is
Illustrative: the two routes as they exist on a retained deployment; latencies are typical rather than measured.

An engineer with a production system on fire is not choosing between discipline and laziness. They are choosing between a thirty-second route and a twenty-minute route while the outage clock runs, and they will choose the fast one every time the emergency is shorter than the code path. Locking the console does not change that calculation; it removes the fast route and leaves the slow one, which means the emergency lasts twenty minutes longer and the next retrospective proposes a break-glass role that reopens the console. Drift is what accumulates in the gap between the two routes, and the width of the gap is the apply latency.

Drift half-life

The number I use to see this is a drift half-life: the time after a clean apply until half of a workspace's resources show a difference in a scheduled plan. Measuring it is cheap. Run plan on a schedule, hourly or daily, against every workspace, record the count of resources with a diff, and watch the curve after each clean apply. Most tools that manage Terraform runs have a drift detection feature that produces the raw data, and a spreadsheet does the rest.

The half-life is the one number that tells you whether your code path is faster than your incidents. A workspace whose half-life is measured in months is one where the code path wins: people change things through code because it is not meaningfully slower than the alternative. A half-life under a week means people are routing around the code path faster than it can keep up, and the fix is in the pipeline, not in IAM.

Drift half-life as a function of apply latency, from a simple decay model A curve falling from left to right on a log-scaled vertical axis. At an apply latency of five minutes, the modelled half-life is many months; at twenty minutes, weeks; at an hour, days; at four hours, under a day. The model assumes a fixed rate of urgent changes with a distribution of urgency, and that each change whose urgency window is shorter than the apply latency goes through the console. Halve the apply latency and the half-life more than doubles Modelled time until half a workspace has drifted, against pull-request-to-production latency 1 yr 1 mo 1 wk 1 day 1 hr 5 min 20 min 1 hr 4 hr 1 day Apply latency, log scale modelled half-life one-week line weeks at twenty minutes under a day at four hours one week: below this, people are routing around you
Illustrative: a decay model in which urgent changes arrive at a steady rate with a spread of urgency windows, and every change whose window is shorter than the apply latency goes through the console; the curve's shape is the point, not its values.

The model behind the curve is simple enough to state. Urgent changes arrive at some steady rate. Each has an urgency window, the time within which it has to land, and those windows are spread across a wide range, from minutes to days. A change goes through the console when its window is shorter than the apply latency and through code otherwise. The share of changes that drift is therefore the share of windows shorter than the latency, and because the windows are spread over orders of magnitude, halving the latency removes more than half of the drifting changes. That is the whole argument for spending on the pipeline: the returns are better than linear.

What locking the console actually does

The read-only console deserves a fair hearing, because it is the standard advice and it is not useless. It makes drift visible: every change now leaves a trace in code or in a break-glass audit log, and a team that could not previously say how much of its estate was hand-edited can. It also removes the accidental drift, the well-meaning tweak by someone who did not know the resource was managed. Those are real gains.

What it does not do is change the race. The urgent change still has an urgency window, the code path still has its latency, and when the window is shorter, one of three things happens: the break-glass role gets used, which is the console with paperwork; the change is made in a tool outside Terraform's view, a Kubernetes manifest applied by hand, a database parameter set through a client, a DNS record changed at the registrar, which is drift the scheduled plan cannot even see; or the change is not made and the outage runs longer. None of those is the outcome the lock was meant to buy. The half-life measurement still applies after the lock, and if it does not improve, the lock has moved the drift rather than removed it.

Shortening the path

On the retained deployments I run, the work of shortening the path is unglamorous and specific. Plans run on every pull request automatically, so the review has the plan in front of it rather than waiting for one. Workspaces are split so that an urgent change to a firewall rule does not plan against the whole estate, which is what makes the plan slow and the diff noisy. Reviews for a small, well-defined class of changes, a rule, a size, a count, are approved by a policy check rather than a person, so that the four-in-the-morning change does not wait for a timezone. The apply runs from the merge, not from a manual trigger someone has to remember. And the state lock is held for as short a time as the tool allows, because a locked state during an incident is the single most reliable way to make someone open the console.

None of that touches IAM. The console stays available, because during a real emergency it should be, and what changes is that it stops being the faster route for anything but the true emergencies. Then the half-life, measured by the scheduled plans, tells you whether it worked. When the number climbs from days to months, the discipline conversation stops happening, not because people became more disciplined but because the race was won by the route you wanted them to take.

TerraformInfrastructure As CodeCloud
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS