9 June 2025 · 5 min read

Restore time is the only backup metric

Backup success rate is a vanity number. The number that decides whether a company survives is how long a restore takes from nothing, and in every public postmortem it ran long.

Every backup dashboard I have seen reports the same thing: last night's job succeeded, the snapshot exists, the retention policy is in force. All true, all reassuring, and none of it is the number that matters. The number that matters is how long it takes to get from nothing, an empty cloud account, a repository and an offsite copy, to a system that passes its production health check. That number is rarely measured, because measuring it is frightening, and the three most instructive postmortems in the industry are instructive precisely because nobody had measured it before the day it mattered.

Three restores that took longer than anyone planned

GitLab's postmortem of its January 2017 database outage is the best-known. An engineer removed the data directory on the wrong server; the regular backups turned out not to be usable; the restore ran from a staging snapshot, which had to be copied across a slow network link, and the copy alone took about eighteen hours, with roughly six hours of production data lost for good.

Roblox's account of its October 2021 outage describes seventy-three hours of downtime for a service with tens of millions of daily users, most of it spent diagnosing a pathological interaction between a service-discovery system and its storage engine before recovery could begin. Atlassian's post-incident review of April 2022 describes a maintenance script that deleted 883 sites belonging to 775 customers, and a restoration that ran for up to two weeks because the recovery process had been designed for one site at a time, not for hundreds at once.

Time to restore in three public postmortems Horizontal bars on a logarithmic scale of hours: GitLab, January 2017, about 18 hours to copy the database back. Roblox, October 2021, 73 hours to full service. Atlassian, April 2022, up to 14 days, about 336 hours, for the last customers restored. Longer than anyone planned, every time Hours to restoration, log scale, from each company's own account GitLab, 2017 about 18 h Roblox, 2021 73 h Atlassian, 2022 336 h Three causes, one finding: the restore path had never been run at full scale. log scale; the Atlassian figure is the upper end of a range that varied by customer
Source: the GitLab, Roblox and Atlassian postmortems.

Three companies with more engineers than most of my clients have employees, three completely different root causes, and one shared property: the restore path had never been exercised under the conditions of the day, so the time it would take was a guess, and the guess was low. That is the pattern I want a small team to take from the large ones. The failure is not that the backup did not exist. It is that the time to use it was unknown.

The empty-account drill

The drill I run for the systems CustomGlide maintains under retained support is simple to describe and uncomfortable to do. Start from a brand-new cloud account, with nothing in it. Give the person running the drill the repository, the offsite backup, and nothing else: no access to the old account, no colleague on chat, no wiki that is not in the repository. Start a stopwatch. Stop it when the production health check passes on the new account.

The elapsed time is the recovery objective. Not the one in the document; the real one. And every moment during the drill when the person had to look something up outside the repository, ask someone for a credential, or click through a console because the step was not scripted, is a finding, written down as it happens. The drill produces two outputs: a number and a list, and the list is the more valuable of the two, because it is the set of reasons the number is what it is.

The empty-account drill as a sequence with the stopwatch running Three lanes: the person running the drill, the new empty account, and the offsite backup. Steps in order: create the account, apply the infrastructure from the repository, fetch the backup, restore the database, deploy the application, and run the health check. Three points are marked where humans usually get involved: a credential that is not in the repository, a console step that is not scripted, and a backup that needs a key nobody can find. Person, stopwatch Empty account Offsite backup apply infrastructure from the repo finding: a credential not in the repo fetch the backup finding: single-holder key restore, then deploy finding: a console step nobody scripted health check passes: stop the clock the number is the objective; the findings are why it is that number
Illustrative: the drill as I run it for a PostgreSQL-backed containerised service; the three findings are the ones that appear most often.

What the drill finds

The findings repeat across systems with a regularity that is almost funny. The backup is encrypted, and the key is in a password manager that belongs to one person, who is on a flight. The infrastructure is in code, but one resource, usually a DNS record or an OAuth client, was created by hand and the code does not know about it. The database restores, but the application's first migration on a fresh database fails because the migrations were written against a schema that already existed. The container image is in a registry in the old account. The health check depends on a third-party service whose credential was rotated last year and stored nowhere.

None of these is exotic. Each is a few minutes to fix once found and a few hours to find during an incident, and the drill converts the second into the first. The list gets shorter each time the drill is run, and the number comes down with it, which is the only way I know to make a recovery objective true rather than aspirational.

Assumed and measured recovery time across repeated drills Two lines over five drills. The assumed recovery time, from the document, stays flat at two hours. The measured time starts far above it, near fourteen hours on the first drill, and falls with each drill as findings are fixed, reaching a little above the assumption by the fifth. The gap on the first drill is the part that would have been discovered during an incident. The document says two hours; the stopwatch disagrees, then converges Hours from empty account to passing health check; illustrative measured assumed 16 12 8 4 0 drill 1 2 3 4 5 14 h 3 h Each drill fixes its findings before the next
Illustrative: a series drawn to show the shape of convergence, not measured results from any client's system.

What to do with the number

The measured number goes in two places. It goes in the runbook, as the current recovery time with the date it was measured, replacing whatever aspiration was there. And it goes in the conversation with whoever owns the business, as a plain statement: if we lost the account tonight, we would be back in this many hours, and here is the list of things that make it that many rather than fewer. That conversation is where the drill earns its cost, because it converts an abstract worry into a priced list of fixes, and most of the fixes are cheap.

It also sets the cadence. A system whose measured recovery is under an hour and whose findings list is empty can be drilled quarterly. A system whose measured recovery is a day and whose list is long is drilled monthly until the list is short, because each drill is the fastest way to shorten it. The drill is not a compliance exercise to be done once and filed. It is how the number gets smaller.

Why the drill has to start from nothing

Restoring into the existing account is a different and easier exercise, and it is the one most teams do when they do anything. It reuses the credentials, the DNS, the registry, the hand-made resources and the person's memory of where things are. It tests the backup file. It does not test the recovery, because the recovery, on the day it matters, may not have the account: a compromised root credential, a billing dispute, a region gone, a script that deleted the wrong thing at scale. The three postmortems above are, in different ways, all stories about the day the surrounding environment could not be relied on, and a drill that relies on it measures the wrong thing.

Starting from nothing is also what makes the drill honest about people. The drill is run by someone who does not carry the system in their head, or by the person who does, with the rule that nothing in their head counts. Every fact they needed and did not find in the repository is a fact that would have been unavailable if they had been the one on the flight.

The metric, restated

Report one number for backups: the hours from an empty account to a passing health check, as last measured, with the date of the measurement. If the number has never been measured, report that, in those words, because it is the truth and it is more useful than a green tick next to last night's job. Then run the drill, fix the list, and run it again. Backup success rate tells you the file exists. Restore time tells you whether the company does.

Disaster RecoveryBackupsPostmortems
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS