Restore time is the only backup metric
Backup success rate is a vanity number. The number that decides whether a company survives is how long a restore takes from nothing, and in every public postmortem it ran long.
Every backup dashboard I have seen reports the same thing: last night's job succeeded, the snapshot exists, the retention policy is in force. All true, all reassuring, and none of it is the number that matters. The number that matters is how long it takes to get from nothing, an empty cloud account, a repository and an offsite copy, to a system that passes its production health check. That number is rarely measured, because measuring it is frightening, and the three most instructive postmortems in the industry are instructive precisely because nobody had measured it before the day it mattered.
Three restores that took longer than anyone planned
GitLab's postmortem of its January 2017 database outage is the best-known. An engineer removed the data directory on the wrong server; the regular backups turned out not to be usable; the restore ran from a staging snapshot, which had to be copied across a slow network link, and the copy alone took about eighteen hours, with roughly six hours of production data lost for good.
Roblox's account of its October 2021 outage describes seventy-three hours of downtime for a service with tens of millions of daily users, most of it spent diagnosing a pathological interaction between a service-discovery system and its storage engine before recovery could begin. Atlassian's post-incident review of April 2022 describes a maintenance script that deleted 883 sites belonging to 775 customers, and a restoration that ran for up to two weeks because the recovery process had been designed for one site at a time, not for hundreds at once.
Three companies with more engineers than most of my clients have employees, three completely different root causes, and one shared property: the restore path had never been exercised under the conditions of the day, so the time it would take was a guess, and the guess was low. That is the pattern I want a small team to take from the large ones. The failure is not that the backup did not exist. It is that the time to use it was unknown.
The empty-account drill
The drill I run for the systems CustomGlide maintains under retained support is simple to describe and uncomfortable to do. Start from a brand-new cloud account, with nothing in it. Give the person running the drill the repository, the offsite backup, and nothing else: no access to the old account, no colleague on chat, no wiki that is not in the repository. Start a stopwatch. Stop it when the production health check passes on the new account.
The elapsed time is the recovery objective. Not the one in the document; the real one. And every moment during the drill when the person had to look something up outside the repository, ask someone for a credential, or click through a console because the step was not scripted, is a finding, written down as it happens. The drill produces two outputs: a number and a list, and the list is the more valuable of the two, because it is the set of reasons the number is what it is.
What the drill finds
The findings repeat across systems with a regularity that is almost funny. The backup is encrypted, and the key is in a password manager that belongs to one person, who is on a flight. The infrastructure is in code, but one resource, usually a DNS record or an OAuth client, was created by hand and the code does not know about it. The database restores, but the application's first migration on a fresh database fails because the migrations were written against a schema that already existed. The container image is in a registry in the old account. The health check depends on a third-party service whose credential was rotated last year and stored nowhere.
None of these is exotic. Each is a few minutes to fix once found and a few hours to find during an incident, and the drill converts the second into the first. The list gets shorter each time the drill is run, and the number comes down with it, which is the only way I know to make a recovery objective true rather than aspirational.
What to do with the number
The measured number goes in two places. It goes in the runbook, as the current recovery time with the date it was measured, replacing whatever aspiration was there. And it goes in the conversation with whoever owns the business, as a plain statement: if we lost the account tonight, we would be back in this many hours, and here is the list of things that make it that many rather than fewer. That conversation is where the drill earns its cost, because it converts an abstract worry into a priced list of fixes, and most of the fixes are cheap.
It also sets the cadence. A system whose measured recovery is under an hour and whose findings list is empty can be drilled quarterly. A system whose measured recovery is a day and whose list is long is drilled monthly until the list is short, because each drill is the fastest way to shorten it. The drill is not a compliance exercise to be done once and filed. It is how the number gets smaller.
Why the drill has to start from nothing
Restoring into the existing account is a different and easier exercise, and it is the one most teams do when they do anything. It reuses the credentials, the DNS, the registry, the hand-made resources and the person's memory of where things are. It tests the backup file. It does not test the recovery, because the recovery, on the day it matters, may not have the account: a compromised root credential, a billing dispute, a region gone, a script that deleted the wrong thing at scale. The three postmortems above are, in different ways, all stories about the day the surrounding environment could not be relied on, and a drill that relies on it measures the wrong thing.
Starting from nothing is also what makes the drill honest about people. The drill is run by someone who does not carry the system in their head, or by the person who does, with the rule that nothing in their head counts. Every fact they needed and did not find in the repository is a fact that would have been unavailable if they had been the one on the flight.
The metric, restated
Report one number for backups: the hours from an empty account to a passing health check, as last measured, with the date of the measurement. If the number has never been measured, report that, in those words, because it is the truth and it is more useful than a green tick next to last night's job. Then run the drill, fix the list, and run it again. Backup success rate tells you the file exists. Restore time tells you whether the company does.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS