13 May 2026 · 5 min read

Price the outage, not the hour

For infrastructure and security work the hour is the wrong unit. The client is buying a lower probability of an expensive failure, which is what an insurer sells, with arithmetic.

The first quotes I wrote for infrastructure work were in hours, because that is how engineering is bought, and every one of them was wrong in the same direction. The client was not buying my hours. They were buying a lower chance of the thing they were afraid of: the gate that stops validating tickets on the busiest day of the year, the database that cannot be restored, the credential that walks out of the building. Hours are what it costs me to reduce that chance. They say nothing about what the reduction is worth, and the gap between the two is where a quote is either fair or foolish.

What the client is actually buying

Uptime Institute's annual survey makes the shape of the thing plain. In its 2026 outage analysis, 57 percent of respondents said their most recent significant outage cost more than 100,000 dollars, and one in five said it cost more than a million, the second year running at that level; the 2025 edition had 54 percent above 100,000. Those are large operators, and a small client's numbers are smaller, but the structure is the same: an outage has a cost, the cost is large relative to the work that prevents it, and the work changes a probability rather than delivering a thing.

Cost of the most recent significant outage, share of respondents Two pairs of horizontal bars. Over 100,000 dollars: 54 percent in the 2024 survey, 57 percent in the 2025 survey. Over one million dollars: about 20 percent in both. What an outage costs the people who had one Share of respondents, Uptime Institute annual surveys Over $100,000, 2024 survey 54% Over $100,000, 2025 survey 57% Over $1 million, 2024 survey one in five Over $1 million, 2025 survey one in five the million-dollar rows are drawn at 20 per cent, as the reports state them
Source: Uptime Institute press releases for the 2025 and 2026 annual outage analyses.

An engagement that reduces a probability is an insurance product, whatever the invoice calls it. And insurers do not price by the hour. They price by the expected loss they are absorbing, and they add a margin for the risk they cannot measure. That arithmetic transfers to infrastructure and security work almost unchanged, and it produces quotes that survive negotiation because every number in them is the client's own.

The premium quote

The quote has four inputs, and all four come from the client. The cost of the failure: what a day without the gate, or a lost database, or a leaked key would cost, in money, using the client's own figures for revenue, penalties and recovery. The baseline probability: how likely that failure is over the retention period, before the work. The post-work probability: how likely it is after. And the retention period itself: how long the client is buying protection for, which is usually the length of the support contract.

The fee is a stated share of the expected loss avoided: the cost of failure, times the reduction in probability, over the period. To that I attach a deductible, which is the residual risk the client keeps, stated plainly: the failures the work does not prevent, and the probability that remains.

The premium quote as a flow Three inputs on the left: cost of failure, baseline probability, and probability after the work. They combine into expected loss avoided over the retention period. A share of that becomes the fee. What remains, the residual probability times the cost, is stated as the deductible the client keeps. Cost of failure Probability before Probability after Expected loss avoided cost x reduction, per period Fee: a stated share of the loss avoided Deductible residual risk, in writing every input is the client's own number; the share is the only one that is mine
Illustrative: the structure of the quote; the share and the inputs vary by engagement.

The same hours, three prices

Here is why the hour was the wrong unit, in numbers I have made up to show the shape. Take forty hours of the same work: a restore drill, a lock-timeout policy for migrations, secrets moved to short-lived tokens, and a runbook. Do it for a hobby site whose worst day costs its owner a few hundred rupees of embarrassment. Do it for a visitor attraction whose gates fail on a festival weekend and refund a day's tickets. Do it for a payments backend whose outage triggers contractual penalties and a regulator's letter. The hours are identical. The loss avoided differs by orders of magnitude, and a fee that is a share of the loss avoided differs with it.

The same forty hours priced for three clients Horizontal bars on a logarithmic scale of expected loss avoided: a hobby site at about 1, a ticketing gate at about 100, a payments backend at about 10,000, in relative units. The fee follows the bar, not the hours. Identical work, different worth Expected loss avoided by the same 40 hours, relative units, log scale Hobby site 1 Ticketing gate 100 Payments backend 10,000 An hourly quote charges all three the same. A premium quote charges each a share of its own bar, and is cheap for the first and fair for the third. made-up magnitudes; the point is the spread, not the values
Illustrative: three client profiles with invented magnitudes, drawn to show why the hour cannot be the unit.

The hobby site should probably not buy the work at all, and the premium quote says so honestly: the loss avoided is too small to be worth a fee that covers my time, and the right answer is a checklist they can follow themselves. The payments backend should buy it and should pay a great deal more than forty hours' worth, and will, because the quote is a fraction of a number their own finance team produced. The ticketing gate sits in between, and the deductible matters most there, because the work reduces the chance of a bad festival weekend and does not eliminate it, and the client should know which failures are still theirs.

Estimating the probabilities without pretending

The two probabilities are the inputs people distrust, and the honest answer is that they are estimates and should be presented as ranges. The baseline comes from three places. Public base rates, such as the outage surveys above, give an order of magnitude for a class of failure. The client's own history gives a second: how many times in the last three years the gate went down, the database needed a restore, a credential had to be rotated in a hurry. And a short review of the system gives a third, because a database with no tested restore has a higher probability of an unrecoverable failure than one with a monthly drill, and that difference is visible in an afternoon.

The post-work probability is the harder one, and I state it as a claim about mechanism rather than a number pulled from the air. A restore drill that runs monthly does not make data loss impossible; it makes an unrecoverable loss depend on two failures in the same month instead of one. A lock timeout on migrations does not prevent bad migrations; it turns a table-wide outage into a failed deployment. Each mechanism converts one kind of failure into a smaller kind, and the reduction in probability is the reduction in the kinds that remain. Written that way, the client can argue with the mechanism, which is a better argument than one about a decimal.

Where the formula stops working

The arithmetic needs a cost of failure, and some work has none that anyone will write down. A refactor that makes the code nicer, a migration between two adequate frameworks, a dashboard that nobody's decision depends on. For those the premium quote produces zero, which is the formula telling you the work should be priced by the hour, if at all, because there is no risk being transferred. That is a useful answer, not a failure of the method.

It also needs honest probabilities, and both sides have reasons to shade them. The client would like the baseline to be low, so the work looks unnecessary, and then high, so the fee looks steep. I would like the reduction to be large. The defence is to use public base rates where they exist, the outage surveys, the secrets-sprawl figures, the migration-incident post-mortems, and to write the assumptions into the quote so that they can be argued about before the work rather than after the failure.

What changed

Quoting this way changed two things about CustomGlide's retained infrastructure work. The conversations moved from "how many hours" to "what does the bad day cost you", which is a conversation the client's leadership can join and the engineering manager could not have held alone. And the deductible line, the plain statement of what is still their risk, became the part of the proposal clients read most carefully, because it is the only part that tells them what they are not buying.

Value-based pricing is old advice, and it usually arrives without arithmetic, as an exhortation to charge what the work is worth. Underwriting is the arithmetic. Cost of failure, probability before, probability after, a share of the difference, and a deductible in writing. It is not a way to charge more. It is a way to charge the right amount, which for some clients is less and for a few is a great deal more, and to be able to say why.

PricingConsultingReliability
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS