Stale is a feature
Most origins that fall over could have kept serving for hours if the CDN had been told a slightly old answer was acceptable. Write that tolerance down per route, in one header.
Most outages I have watched from the outside had the same shape: the origin stopped answering, the CDN in front of it had a perfectly good copy of every page from a few minutes earlier, and the CDN served error pages anyway, because nobody had told it that a slightly old answer was acceptable. The HTTP specification has had a way to say that since 2010. Almost nobody uses it, and the reason is not ignorance of the header but the absence of a number to put in it. The number is the staleness budget, and it is different for every route.
The unused half of the specification
RFC 5861 added two extensions to Cache-Control. stale-while-revalidate lets a cache serve a stale copy immediately while fetching a fresh one in the background, which is a latency feature. stale-if-error lets a cache serve a stale copy when the origin returns an error or cannot be reached, for a stated number of seconds past expiry, which is an availability feature. The second is the one that turns a dead origin into a slow news day, and it is the rarer of the two by an order of magnitude.
The HTTP Archive's Web Almanac has counted directive usage across millions of pages. In the 2019 chapter, stale-while-revalidate appeared on 2.4 percent of responses and stale-if-error on 0.2 percent. In 2020 the figures were 2.2 and 0.2 percent on mobile, and in 2021 2.4 and 0.2. Fewer than one response in forty carries the latency directive; one in five hundred carries the outage one; and the numbers did not move across three years while the CDN industry grew around them.
What the header changes
Consider a request for a page during an origin failure. Without the directive, the CDN's copy expires, the CDN asks the origin, the origin returns a 503 or nothing, and the CDN passes the failure to the user. With stale-if-error set to, say, 86,400, the CDN's copy expires, the CDN asks the origin, the origin fails, and the CDN serves the copy it already has, for up to a day past its expiry, while retrying the origin in the background. The user sees a page that is a few minutes old. The operator sees an alert instead of a customer.
Two cautions belong here. The first is that CDNs differ in whether and how they honour stale-if-error; some support it as written, some implement the same idea under a product setting with a different name, and some ignore it. Cloudflare's Cache-Control documentation is where I check for the sites I run on it, and every other CDN has an equivalent page. The second is that the header only helps for responses the CDN was allowed to cache in the first place. A route marked no-store has no copy to fall back on, and that is often correct, which brings us to the number.
The staleness budget
The reason the directive goes unused is that it demands a number, and the number is a product decision disguised as a configuration value. It is the answer to: for this route, for how many seconds is a wrong answer cheaper than no answer? I call it the staleness budget, and it is different for every kind of route in a way that no site-wide default can capture.
The poles are easy. A static portfolio, a lawyer's site, a game studio's landing page: nothing on them changes in a way that a visitor during an outage would notice, and a budget of a day is conservative. Anyone who visited during the hour the origin was down saw exactly what they would have seen the day before, which is the whole site. At the other pole, a ticket scan-in endpoint has a budget of zero, because a stale answer to "has this ticket already been used" is a door held open for a duplicate, and the honest behaviour when the origin is down is to refuse and say why. An account balance is the same. Between the poles is where the decision lives, and it is a decision: a product listing can be an hour old because the checkout will check stock again, a pricing page can be ten minutes old but not a day because an old price shown is a price someone will expect to pay, a dashboard can be a minute old if it shows its own timestamp so that the reader knows.
Three objections
The first objection is that stale data is wrong data, and a site should not knowingly serve wrong data. The budget is the answer, because it makes the wrongness bounded and explicit: a page at most ten minutes old is wrong in a way the product has agreed to, while an error page is wrong in a way nobody agreed to and that carries no information at all. The routes where any staleness is unacceptable get a budget of zero and are excluded, which is the objection taken seriously rather than dismissed.
The second is that the origin does not go down. It does, and not only in incidents. Deploys restart processes, certificates expire, a database failover takes the application with it for a minute, a dependency's outage becomes a 500 on every page that calls it. stale-if-error covers all of those, because it triggers on any error response or an unreachable origin, and a one-minute deploy blip on a site with a one-hour budget is invisible to visitors. Most CDNs expose the same idea under a product name, often something like serving stale on origin error, and the header is the portable way to ask for it.
The third is subtler: purging. A cache purge deletes the copy, and a deleted copy cannot be served stale, so a team that purges the whole cache on every deploy has removed its own safety net at the moment the origin is most likely to be restarting. The fix is to purge by key rather than wholesale, invalidating only the routes whose content changed, and to sequence deploys so that the purge happens after the new origin is healthy rather than before. The budget only protects a copy that still exists.
Writing it down
The practice that follows is small. Every route in the site or the API gets a budget, in seconds, recorded next to its cache policy, and the budget is emitted as stale-if-error on that route's responses. Routes with a budget of zero get no directive and, ideally, an explicit error page that says the service is unavailable rather than a generic one. The budget list becomes the document that describes the site's outage tolerance, and it is the only such document most sites will ever have, because it is one line per route and it is enforced by the CDN rather than by a runbook.
The same number also settles the stale-while-revalidate question. If a route can tolerate being a minute old during an outage, it can tolerate being a minute old during a background refresh, so the latency directive gets the same value or a smaller one, and the two are set together. The web's numbers say that most sites set neither. The staleness budget is how you decide what to set, and once decided, it takes one header to make the CDN keep the lights on while you fix the origin.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS