24 April 2025 · 5 min read

Publish the denominator

Impact reports count outputs and drop two numbers: how many were eligible, and what would have happened anyway. A metric with no baseline would not ship, so do not fund one.

An impact report lands in the inbox with a number on the cover: twelve thousand meals served, four hundred students trained, three hundred families reached. The number is true, it is large, and it tells the reader almost nothing, because two other numbers are missing. How many people were eligible, so that the reach can be read as a fraction? And what would have happened to those people without the programme, so that the outcome can be read as a change? I hold the software I run to a standard where a metric without a denominator and a baseline is not a metric; it is a log line. The giving I do should meet the same standard, and the rule that makes it meet it is one sentence long.

The denominator rule

Every impact figure is published as a numerator over the eligible population, with the source of the counterfactual stated, or it is labelled an output rather than an outcome. That is the whole rule. Twelve thousand meals is an output. Twelve thousand meals to a district where forty thousand children are eligible, which is thirty percent, is a reach figure with its denominator shown. A ten-point rise in attendance among the children fed, against a control group whose attendance did not move, is an outcome with its counterfactual named. Each of those is a legitimate thing to report. Only the last is impact, and the report should say which of the three it is presenting.

The ladder from output to outcome to counterfactual, with the denominator at each rung Three rungs drawn as boxes rising left to right. Output: what was delivered, with no denominator; twelve thousand meals. Reach: delivered over eligible; twelve thousand of forty thousand, thirty percent. Outcome against a counterfactual: the change in the people reached minus the change in a comparison group; attendance up ten points against no change. The top rung is highlighted as the only one that is impact. Three rungs, and only the top one is impact Output 12,000 meals served no denominator Reach 12,000 of 40,000 eligible denominator shown: 30 percent Outcome, counterfactual attendance up 10 points vs a group that stayed Each rung is honest to report. The report has to say which rung it is on.
Illustrative: the ladder as the rule uses it; the numbers are an example.

The rule is not a demand that every programme run a randomised trial. It is a demand for labelling. A report that says "we served twelve thousand meals; we do not know the eligible population and we have no comparison group" is complying with the rule, and it is more useful than one that says "twelve thousand lives changed", because the reader can tell what has been measured and what has been assumed.

What a real outcome number looks like

The reason to insist on the top rung is that when it is measured, the numbers are often smaller and more informative than the outputs suggest. The clearest example I know from Indian education is the evaluation of Mindspark, a computer-assisted learning programme, by Muralidharan, Singh and Ganimian, published in the American Economic Review in 2019. Access was allocated by lottery, which supplies the counterfactual, and after four and a half months the winners scored 0.37 standard deviations higher in maths and 0.23 higher in Hindi than the losers; the authors' estimate for ninety days of actual attendance is 0.6 and 0.39 standard deviations. Those are large effects by the standards of the education literature, and they are stated as a difference against a group that did not get the programme, which is the only way an effect can be stated at all.

Measured learning gains from the Mindspark evaluation, in standard deviations Grouped horizontal bars. Lottery winners after 4.5 months: maths 0.37, Hindi 0.23. Estimate for 90 days of attendance: maths 0.60, Hindi 0.39. Maths bars are gold and Hindi bars are blue. An outcome with its counterfactual named Test score gain over the lottery losers, standard deviations maths Hindi Winners, 4.5 months 0.37 0.23 Per 90 days attended 0.60 0.39 Six hundred pixels per standard deviation; the second pair is the instrumental estimate. the denominator is the lottery: everyone who applied, winners and losers alike
Source: Muralidharan, Singh and Ganimian, Disrupting Education? Experimental Evidence on Technology-Aided Instruction in India, American Economic Review, 2019.

The same discipline is what let Teaching at the Right Level move from a small trial to state-wide programmes: J-PAL's account of that work reports, among other results, that learning camps in Uttar Pradesh doubled the share of children who could read a paragraph or a story, and the doubling is a comparison, not a count. Nobody funded those programmes on the strength of the number of children who attended.

The three-line template

The rule becomes a template that fits on an index card and that I now ask for before giving. Line one: the output, as delivered, with its unit. Line two: the eligible population and its source, so that line one can be divided by it. Line three: the counterfactual, which is one of four things: a randomised comparison group, a matched comparison group, a before-and-after measurement with the caveat stated, or the words "none; this figure is an output". A report that fills in all three lines has told the reader everything they need to weigh it, whichever rung it is on.

One programme reported three ways A table with three rows. As reach: 400 students trained. As a fraction of eligible: 400 of 6,000 eligible in the district, 6.7 percent, source the district enrolment register. Against a counterfactual: completion of the course 82 percent, against 79 percent for a matched comparison group of non-participants, a 3-point difference with the caveat that the groups differ in motivation. The third row is highlighted. The same programme, on each rung REPORTED AS THE FIGURE WHAT IT LETS YOU JUDGE Output 400 students trained that work was done Reach 400 of 6,000 eligible, 6.7% source: enrolment register how much of the problem the programme touched Outcome 82% employed at six months, against 79% in a matched group caveat: groups differ in motivation whether the programme changed anything: 3 points Only the last row is a claim about impact, and it is the smallest number on the page.
Illustrative: an invented training programme reported on each rung, to show what each line of the template adds.

The template's third row is where reports get uncomfortable, because the honest number is usually smaller than the headline. Eighty-two percent employed sounds like a result until the comparison group is at seventy-nine, at which point the programme's contribution is three points and a caveat. That is not a reason to hide the row. It is the reason the row exists: a donor who knows the contribution is three points can ask whether three points for the cost is a good use of money, and that question is the one philanthropy is supposed to be answering.

Asking for the template is not an adversarial act, and the way to ask matters. I send it before giving, with the offer that if the second or third line is missing because measuring it costs money, I will fund the measurement as part of the gift. Most organisations doing real work already know their eligible population roughly and have a before-and-after figure they were unsure whether to show; what they lack is a donor who would rather see the small honest number than the large vague one. Saying so in advance changes what comes back.

The limit: unknowable denominators

The rule has a limit, and the limit needs its own line rather than silence. Some denominators cannot be counted. The number of children in a district who would benefit from a reading programme is not in any register; the number of families eligible for a relief effort during a flood is a guess made while the water is rising. The rule for those cases is that the denominator is estimated with a range and the method stated, never omitted. "Between five and eight thousand eligible, estimated from census population and the enrolment ratio" is a denominator. "Thousands" is not. A range is honest about uncertainty in a way that a missing number is not, and a reader can carry the range through the arithmetic to see what the reach figure could be at either end.

Counterfactuals have the same limit. There are programmes for which no comparison group can ethically exist, and the honest report says so and falls back to before-and-after with the caveat that other things changed too. The rule does not forbid that report. It forbids the report that presents before-and-after as if it were a comparison.

Why a technologist should insist

I am strict about this because I know what the alternative looks like from the other side. In production, a metric without a denominator is a count that goes up and to the right whatever is happening: requests served, not the share of requests that succeeded; users signed up, not the share who came back. A count without a baseline is an alert that fires on Monday morning because Monday is busier than Sunday. Nobody who has run a service accepts those numbers as evidence of health, and the observability practice I follow exists precisely to replace counts with rates and rates with comparisons.

Giving is not different in kind. A programme is a system whose purpose is to change something in a population, and the question of whether it did is a measurement question with a denominator and a baseline, like every other measurement question. The denominator rule is the observability standard applied to the one place where people are most tempted to relax it, because the numbers are about kindness rather than uptime. The kindness is real. The number should be too.

PhilanthropyMeasurementEvidence
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS