9 January 2025 · 5 min read

Stale is a feature

Most origins that fall over could have kept serving for hours if the CDN had been told a slightly old answer was acceptable. Write that tolerance down per route, in one header.

Most outages I have watched from the outside had the same shape: the origin stopped answering, the CDN in front of it had a perfectly good copy of every page from a few minutes earlier, and the CDN served error pages anyway, because nobody had told it that a slightly old answer was acceptable. The HTTP specification has had a way to say that since 2010. Almost nobody uses it, and the reason is not ignorance of the header but the absence of a number to put in it. The number is the staleness budget, and it is different for every route.

The unused half of the specification

RFC 5861 added two extensions to Cache-Control. stale-while-revalidate lets a cache serve a stale copy immediately while fetching a fresh one in the background, which is a latency feature. stale-if-error lets a cache serve a stale copy when the origin returns an error or cannot be reached, for a stated number of seconds past expiry, which is an availability feature. The second is the one that turns a dead origin into a slow news day, and it is the rarer of the two by an order of magnitude.

The HTTP Archive's Web Almanac has counted directive usage across millions of pages. In the 2019 chapter, stale-while-revalidate appeared on 2.4 percent of responses and stale-if-error on 0.2 percent. In 2020 the figures were 2.2 and 0.2 percent on mobile, and in 2021 2.4 and 0.2. Fewer than one response in forty carries the latency directive; one in five hundred carries the outage one; and the numbers did not move across three years while the CDN industry grew around them.

Share of web responses carrying the two stale directives, Web Almanac 2019 to 2021 Grouped horizontal bars for three years. stale-while-revalidate: 2.4 percent in 2019, 2.2 in 2020, 2.4 in 2021. stale-if-error: 0.2 percent in each year. Both series are flat and small. The outage directive is used on one response in five hundred Percent of responses carrying the directive, mobile crawl stale-while-revalidate stale-if-error 2019 2.4 0.2 2020 2.2 0.2 2021 2.4 0.2 Scale: 100 pixels per percentage point. Both lines are flat across the three editions.
Source: HTTP Archive Web Almanac caching chapters for 2019, 2020 and 2021; mobile figures where the chapter splits them.

What the header changes

Consider a request for a page during an origin failure. Without the directive, the CDN's copy expires, the CDN asks the origin, the origin returns a 503 or nothing, and the CDN passes the failure to the user. With stale-if-error set to, say, 86,400, the CDN's copy expires, the CDN asks the origin, the origin fails, and the CDN serves the copy it already has, for up to a day past its expiry, while retrying the origin in the background. The user sees a page that is a few minutes old. The operator sees an alert instead of a customer.

A request during an origin failure, with and without stale-if-error Two rows of steps. Without the directive: cached copy expires, CDN asks origin, origin fails, user receives an error page. With the directive: cached copy expires, CDN asks origin, origin fails, user receives the stale copy while the CDN retries in the background. The two rows differ only at the last step. Same failure, different last step WITHOUT copy expires CDN asks origin origin fails user gets an error WITH stale-if-error copy expires CDN asks origin origin fails user gets the copy CDN retries behind it Three identical steps; the header decides only what happens when the origin says no.
Illustrative: the request path during an origin failure; whether a given CDN honours the directive, and for which status codes, is a matter for its documentation.

Two cautions belong here. The first is that CDNs differ in whether and how they honour stale-if-error; some support it as written, some implement the same idea under a product setting with a different name, and some ignore it. Cloudflare's Cache-Control documentation is where I check for the sites I run on it, and every other CDN has an equivalent page. The second is that the header only helps for responses the CDN was allowed to cache in the first place. A route marked no-store has no copy to fall back on, and that is often correct, which brings us to the number.

The staleness budget

The reason the directive goes unused is that it demands a number, and the number is a product decision disguised as a configuration value. It is the answer to: for this route, for how many seconds is a wrong answer cheaper than no answer? I call it the staleness budget, and it is different for every kind of route in a way that no site-wide default can capture.

Staleness budgets by route type, from a day to zero A table of route types and budgets in seconds. A marketing or portfolio page: 86,400, a day. A blog post: 86,400. A product listing: 3,600, an hour. A pricing page: 600, ten minutes. A dashboard: 60, a minute. A ticket validation endpoint: 0. An account balance: 0. The rows with a non-zero budget can carry stale-if-error; the zero rows must fail honestly. How long is a wrong answer cheaper than no answer? ROUTE BUDGET WHY Marketing or portfolio page a day nothing changes by the hour Blog post a day a typo fix can wait Product listing an hour checkout re-checks stock Pricing page ten minutes an old price is a promise Dashboard a minute shown with its timestamp Ticket validation, balances zero a stale yes is a real loss Every non-zero row can carry stale-if-error. The zero rows must fail, and say so.
Illustrative: budgets I would set for these route types; the values are judgements, and the point is that each route gets its own.

The poles are easy. A static portfolio, a lawyer's site, a game studio's landing page: nothing on them changes in a way that a visitor during an outage would notice, and a budget of a day is conservative. Anyone who visited during the hour the origin was down saw exactly what they would have seen the day before, which is the whole site. At the other pole, a ticket scan-in endpoint has a budget of zero, because a stale answer to "has this ticket already been used" is a door held open for a duplicate, and the honest behaviour when the origin is down is to refuse and say why. An account balance is the same. Between the poles is where the decision lives, and it is a decision: a product listing can be an hour old because the checkout will check stock again, a pricing page can be ten minutes old but not a day because an old price shown is a price someone will expect to pay, a dashboard can be a minute old if it shows its own timestamp so that the reader knows.

Three objections

The first objection is that stale data is wrong data, and a site should not knowingly serve wrong data. The budget is the answer, because it makes the wrongness bounded and explicit: a page at most ten minutes old is wrong in a way the product has agreed to, while an error page is wrong in a way nobody agreed to and that carries no information at all. The routes where any staleness is unacceptable get a budget of zero and are excluded, which is the objection taken seriously rather than dismissed.

The second is that the origin does not go down. It does, and not only in incidents. Deploys restart processes, certificates expire, a database failover takes the application with it for a minute, a dependency's outage becomes a 500 on every page that calls it. stale-if-error covers all of those, because it triggers on any error response or an unreachable origin, and a one-minute deploy blip on a site with a one-hour budget is invisible to visitors. Most CDNs expose the same idea under a product name, often something like serving stale on origin error, and the header is the portable way to ask for it.

The third is subtler: purging. A cache purge deletes the copy, and a deleted copy cannot be served stale, so a team that purges the whole cache on every deploy has removed its own safety net at the moment the origin is most likely to be restarting. The fix is to purge by key rather than wholesale, invalidating only the routes whose content changed, and to sequence deploys so that the purge happens after the new origin is healthy rather than before. The budget only protects a copy that still exists.

Writing it down

The practice that follows is small. Every route in the site or the API gets a budget, in seconds, recorded next to its cache policy, and the budget is emitted as stale-if-error on that route's responses. Routes with a budget of zero get no directive and, ideally, an explicit error page that says the service is unavailable rather than a generic one. The budget list becomes the document that describes the site's outage tolerance, and it is the only such document most sites will ever have, because it is one line per route and it is enforced by the CDN rather than by a runbook.

The same number also settles the stale-while-revalidate question. If a route can tolerate being a minute old during an outage, it can tolerate being a minute old during a background refresh, so the latency directive gets the same value or a smaller one, and the two are set together. The web's numbers say that most sites set neither. The staleness budget is how you decide what to set, and once decided, it takes one header to make the CDN keep the lights on while you fix the origin.

CDNCachingCloudflare
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS