One percent of users cannot tell you much
A staged rollout is a statistical test, and at one percent the sample is too small to see the crash-rate rise Play penalises. Write each halt rule from what the stage can detect.
A staged rollout feels like caution. Ship to one percent, watch the crash rate, widen to five, watch again, and so on to everyone. I run that ladder for eight apps on Google Play, and for a long time I treated the early stages as if they were telling me something about the crash rate. They were not, or not the thing I thought. At one percent of a small app's users, the sample is too small to see the very rise in crashes that Play would penalise, and a stage that cannot see the thing you are watching for is not a test of it. It is a delay with a dashboard.
What Play is watching for
Google Play measures app quality through Android vitals, and its documentation states the thresholds it acts on, with the definitions of a user-perceived crash and ANR in the Android vitals developer guide. An app whose user-perceived crash rate exceeds 1.09 percent of daily users, or whose user-perceived ANR rate exceeds 0.47 percent, is over the bad-behaviour threshold and becomes less discoverable on the store; on any single device model the threshold is 8 percent, and an app over it there may get a warning on its listing for that device. Those are the numbers a rollout is guarding against, so the question for each stage is whether it could detect a rise across them.
The stage as a statistical test
A crash rate is a proportion, and detecting a change in a proportion needs a sample whose size depends on how small the change is. Suppose the baseline is right at the threshold, 1.09 percent, and the new release has quietly doubled it to 2 percent, which would put the app well over the line. How many sessions does a stage need to observe before that doubling shows up as more than noise?
The standard one-sided calculation, at 95 percent confidence and 80 percent power, gives about a thousand sessions to distinguish 2 percent from 1.09. To see a rise to 1.5 percent takes over four thousand. To see 3 percent takes under three hundred, and 5 percent under a hundred. The relationship is steep: halving the size of the rise you want to catch roughly quadruples the sample you need.
Now put a small app on that chart. An app with two thousand daily sessions, at a one percent rollout, collects twenty sessions a day. The stage would need two months to see a rise to 2 percent, and the ladder widens it after two days. So the one percent stage of that app is not a test of whether the crash rate rose to 2 percent. It cannot be. What it can detect in two days, at forty sessions, is a catastrophe: a release that crashes for most users on launch. That is a real thing to guard against, and it is the only thing the stage guards against.
The detectable-defect floor
The rule I use is to compute, for each stage of the ladder, the smallest rise in crash rate that the stage can detect with the sessions it will actually gather before the ladder widens. I call that the detectable-defect floor, and it is arithmetic on three numbers the team already has: the app's daily sessions, the stage's percentage, and the number of days the stage will run. If the floor is above the threshold you care about, the stage is not a test of that threshold, and the halt condition written for it should say what it can see.
Writing the halt condition that way changes what people do with the dashboard. At one percent, nobody stares at a crash-rate graph, because forty sessions cannot draw one; the halt condition is "any crash on launch in more than a handful of sessions", which a person can check in a minute, and the stage's job is to run for two days and confirm the app opens. At fifty percent the crash rate is finally a number with meaning, and the halt condition is a rise the stage can actually see.
The per-device threshold makes it worse
The 8 percent per-device threshold sounds generous and is the one a small rollout is least able to guard. A crash that affects one device model, a particular manufacturer's camera driver, say, shows up as a large rate on that model and a tiny rise in the overall rate, and Play judges the model separately. At one percent of a small app, the number of sessions from any single model is a handful, and a crash rate on a handful of sessions is a coin toss. The stage that could detect a device-specific regression is the one at which each model contributes a few hundred sessions, which for most models on most small apps is the last stage or none.
That argues for a different early check than the crash graph: a list of the device models that have reported at all, compared with the list from the previous release, so that a model that has gone silent, because the app now crashes before it can report, is at least visible as an absence. Absence is a weaker signal than a rate, and it is the only one the early stages can give.
What this does to the early stages
It shortens them, honestly. If a one percent stage can only detect a launch crash, and a launch crash shows up in the first hour, holding the stage for two days is a delay that protects nobody. The ladder for a small app should spend its time in the middle rungs, where the sample grows fast enough to see a real rise before the release reaches everyone, and should move through the bottom rung as soon as the catastrophic check passes.
It also explains a thing every release engineer has felt: the early stages always look fine. Of course they do. A stage that can only detect disasters reports no disasters almost every time, and the absence of a signal from a stage that could not have produced one gets read as reassurance. The floor puts a name on that, and the name stops the reassurance from being counted as evidence.
For an app with hundreds of thousands of daily sessions the arithmetic changes completely: one percent is thousands of sessions a day, the floor at the bottom rung sits near the threshold, and the classic ladder works as advertised. The rule is not that one percent stages are useless. It is that the stage's power is a number, the number comes from the app's size, and the halt condition should be written from the number rather than from the shape of the ladder.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS