30 August 2026 · 5 min read

One percent of users cannot tell you much

A staged rollout is a statistical test, and at one percent the sample is too small to see the crash-rate rise Play penalises. Write each halt rule from what the stage can detect.

A staged rollout feels like caution. Ship to one percent, watch the crash rate, widen to five, watch again, and so on to everyone. I run that ladder for eight apps on Google Play, and for a long time I treated the early stages as if they were telling me something about the crash rate. They were not, or not the thing I thought. At one percent of a small app's users, the sample is too small to see the very rise in crashes that Play would penalise, and a stage that cannot see the thing you are watching for is not a test of it. It is a delay with a dashboard.

What Play is watching for

Google Play measures app quality through Android vitals, and its documentation states the thresholds it acts on, with the definitions of a user-perceived crash and ANR in the Android vitals developer guide. An app whose user-perceived crash rate exceeds 1.09 percent of daily users, or whose user-perceived ANR rate exceeds 0.47 percent, is over the bad-behaviour threshold and becomes less discoverable on the store; on any single device model the threshold is 8 percent, and an app over it there may get a warning on its listing for that device. Those are the numbers a rollout is guarding against, so the question for each stage is whether it could detect a rise across them.

Google Play's bad-behaviour thresholds Three horizontal bars, in per cent of daily users: user-perceived crash rate 1.09, user-perceived ANR rate 0.47, and the per-device threshold for either, 8. The per-device bar is much longer. The lines a rollout must not cross Per cent of daily users, Android vitals bad-behaviour thresholds Crash rate, overall 1.09% ANR rate, overall 0.47% Either, on one device model 8% The overall crash threshold is the one a small rollout stage cannot see a move across.
Source: Play Console Help, Android vitals, and the Android vitals documentation.

The stage as a statistical test

A crash rate is a proportion, and detecting a change in a proportion needs a sample whose size depends on how small the change is. Suppose the baseline is right at the threshold, 1.09 percent, and the new release has quietly doubled it to 2 percent, which would put the app well over the line. How many sessions does a stage need to observe before that doubling shows up as more than noise?

The standard one-sided calculation, at 95 percent confidence and 80 percent power, gives about a thousand sessions to distinguish 2 percent from 1.09. To see a rise to 1.5 percent takes over four thousand. To see 3 percent takes under three hundred, and 5 percent under a hundred. The relationship is steep: halving the size of the rise you want to catch roughly quadruples the sample you need.

Sessions a stage needs to detect a crash-rate rise from the 1.09 percent threshold Four columns: to detect a rise to 1.5 percent, about 4,400 sessions; to 2 percent, about 1,000; to 3 percent, about 270; to 5 percent, about 80. At 95 percent confidence and 80 percent power, one-sided. What a stage has to see before it can see anything Sessions needed, one-sided, 95% confidence, 80% power, baseline 1.09% 4,400 rise to 1.5% 1,000 to 2% 270 to 3% 80 to 5%
Illustrative: computed from the standard normal-approximation sample size for comparing an observed proportion with a fixed baseline; the thresholds are Play's, the rises are examples.

Now put a small app on that chart. An app with two thousand daily sessions, at a one percent rollout, collects twenty sessions a day. The stage would need two months to see a rise to 2 percent, and the ladder widens it after two days. So the one percent stage of that app is not a test of whether the crash rate rose to 2 percent. It cannot be. What it can detect in two days, at forty sessions, is a catastrophe: a release that crashes for most users on launch. That is a real thing to guard against, and it is the only thing the stage guards against.

The detectable-defect floor

The rule I use is to compute, for each stage of the ladder, the smallest rise in crash rate that the stage can detect with the sessions it will actually gather before the ladder widens. I call that the detectable-defect floor, and it is arithmetic on three numbers the team already has: the app's daily sessions, the stage's percentage, and the number of days the stage will run. If the floor is above the threshold you care about, the stage is not a test of that threshold, and the halt condition written for it should say what it can see.

A rollout ladder with a detectable floor and a halt condition on each rung Five rungs from the bottom: 1 percent for two days, floor: only a launch crash for most users; 5 percent for two days, floor: a rise to about 5 percent; 20 percent for three days, floor: a rise to about 2 percent; 50 percent for three days, floor: a rise to about 1.5 percent; 100 percent, ongoing, floor: the threshold itself. Each rung's halt condition is written against its floor. Halt on what the rung can see An example ladder for an app with a few thousand daily sessions STAGE SESSIONS SEEN CAN DETECT, AND SO HALTS ON 100%, ongoing all of them the 1.09% threshold itself 50%, three days about 3,000 a rise to about 1.5% 20%, three days about 1,200 a rise to about 2% 5%, two days about 200 a rise to about 5%, or a crash loop 1%, two days about 40 a crash for most users on launch Sessions assume 2,000 a day; recompute the floors for your own numbers.
Illustrative: one ladder with floors computed from the chart above for an app with about two thousand daily sessions; the structure is the point.

Writing the halt condition that way changes what people do with the dashboard. At one percent, nobody stares at a crash-rate graph, because forty sessions cannot draw one; the halt condition is "any crash on launch in more than a handful of sessions", which a person can check in a minute, and the stage's job is to run for two days and confirm the app opens. At fifty percent the crash rate is finally a number with meaning, and the halt condition is a rise the stage can actually see.

The per-device threshold makes it worse

The 8 percent per-device threshold sounds generous and is the one a small rollout is least able to guard. A crash that affects one device model, a particular manufacturer's camera driver, say, shows up as a large rate on that model and a tiny rise in the overall rate, and Play judges the model separately. At one percent of a small app, the number of sessions from any single model is a handful, and a crash rate on a handful of sessions is a coin toss. The stage that could detect a device-specific regression is the one at which each model contributes a few hundred sessions, which for most models on most small apps is the last stage or none.

That argues for a different early check than the crash graph: a list of the device models that have reported at all, compared with the list from the previous release, so that a model that has gone silent, because the app now crashes before it can report, is at least visible as an absence. Absence is a weaker signal than a rate, and it is the only one the early stages can give.

What this does to the early stages

It shortens them, honestly. If a one percent stage can only detect a launch crash, and a launch crash shows up in the first hour, holding the stage for two days is a delay that protects nobody. The ladder for a small app should spend its time in the middle rungs, where the sample grows fast enough to see a real rise before the release reaches everyone, and should move through the bottom rung as soon as the catastrophic check passes.

It also explains a thing every release engineer has felt: the early stages always look fine. Of course they do. A stage that can only detect disasters reports no disasters almost every time, and the absence of a signal from a stage that could not have produced one gets read as reassurance. The floor puts a name on that, and the name stops the reassurance from being counted as evidence.

For an app with hundreds of thousands of daily sessions the arithmetic changes completely: one percent is thousands of sessions a day, the floor at the bottom rung sits near the threshold, and the classic ladder works as advertised. The rule is not that one percent stages are useless. It is that the stage's power is a number, the number comes from the app's size, and the halt condition should be written from the number rather than from the shape of the ladder.

AndroidStaged RolloutsRelease Engineering
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS