15 May 2026 · 5 min read

The agreement ceiling is not a ceiling

Folk wisdom says a model cannot beat inter-annotator agreement. It can, because agreement compares two noisy people while the model is scored against a majority better than either.

There is a sentence that gets said in every labelling project, usually by someone trying to be realistic: the model cannot do better than the annotators agree with each other. If two people agree on 82 percent of items, then 82 percent is the ceiling, and a model scoring above it is measuring noise. It sounds like humility and it is a mistake. The agreement rate compares two noisy annotators with each other. The model is compared with a majority label that is more accurate than either of them. Those are different numbers, and the second one is higher.

Two numbers that look alike

Take the simplest case: binary labels, independent annotators, each right with the same probability p. Two annotators agree when both are right or both are wrong, so the pairwise agreement a is p squared plus one minus p squared. Invert that and the per-annotator accuracy is one plus the square root of two a minus one, all over two. An agreement rate of 82 percent gives a per-annotator accuracy of 90 percent. The annotators are better than their agreement suggests, because agreement counts every disagreement against both of them.

Now build the gold label by majority vote of five annotators at 90 percent each. The majority is wrong only when three or more of the five are wrong, and with independent errors that happens about 0.9 percent of the time. The gold label is 99 percent accurate. A model is scored against that label, so a model that is 95 percent accurate against the truth scores about 95 percent against the gold set, well above the 82 percent that was supposed to be the ceiling, and it is not measuring noise. It is measuring its accuracy, against a label that is good enough to measure it.

Accuracy of a majority label against the number of annotators Three lines over one, three, five and seven annotators. With annotators 90 percent accurate the majority label rises from 90 to 97.2 to 99.1 to 99.7 percent. At 80 percent it rises from 80 to 89.6 to 94.2 to 96.7. At 70 percent it rises from 70 to 78.4 to 83.7 to 87.4. The majority is better than the people in it Majority-label accuracy, per cent, independent binary annotators p = 0.9 p = 0.8 p = 0.7 100 90 80 70 60 1 3 5 7 Annotators per item, majority label 99.7 96.7 87.4
Illustrative: computed from the binomial probability that a majority of k independent annotators is correct, at three per-annotator accuracies; the independence assumption is discussed below.

What the real datasets show

The pattern shows up in the datasets people actually train on. The SNLI corpus of Bowman et al. (2015) collected five labels for each item in its validation and test sets and reports that three of five annotators agreed on 98 percent of items while all five agreed on only 58.3 percent. Read the unanimous figure as the agreement ceiling and you would conclude that no model can score above the high fifties on SNLI. Models passed that years ago, because the gold label is the majority, and the majority exists on 98 percent of items.

The other direction matters too. A gold set is not perfect just because it was voted on, and its errors put a real ceiling on what a model can appear to do. Northcutt, Athalye and Mueller's survey of test-set label errors found that about 6 percent of the ImageNet validation set carries an incorrect label, 2,916 images. A model that is right about those images is scored as wrong, and a model that has learnt the errors is scored as right. The ceiling exists; it is the accuracy of the gold label, and it is a number you can raise with more annotators per item rather than with more items.

Agreement in SNLI and label errors in ImageNet Two horizontal bars for SNLI: items where at least three of five annotators agreed, 98 percent, and items where all five agreed, 58.3 percent. A third bar for ImageNet's validation set: about 6 percent of labels found to be errors. The majority exists where unanimity does not Per cent of items, two public datasets SNLI, 3 of 5 agree 98 SNLI, all 5 agree 58.3 ImageNet val, label errors about 6 The gold label's accuracy is the ceiling; unanimity is not.
Source: Bowman et al. (2015) for the SNLI agreement figures; Northcutt et al. (2021) for the ImageNet validation error estimate.

The agreement inversion

The practical tool is the inversion in the first section, which I use as a small table when budgeting a labelling job. Measure pairwise agreement on a pilot batch. Invert it to per-annotator accuracy. Then read off the majority-label accuracy for one, three, five and seven annotators per item, and pick the count that puts the gold label's accuracy comfortably above the accuracy you expect the model to reach. If you expect a model to hit 95 percent, a gold set that is 97 percent accurate is barely able to tell you so; at 99 percent it can. That decision fixes the annotators per item, and the budget follows.

The inversion says something else that is easy to miss: the way to raise the ceiling is more annotators per item, not more items. Ten thousand items with one label each have a ceiling equal to one annotator's accuracy, whatever that is, and no amount of additional single-labelled items moves it. Two thousand items with five labels each have a ceiling near 99 percent and enough items to measure most things. For evaluation sets in particular, the second design is almost always the right one.

Which pair each number compares Three boxes: annotator, majority label, and model. An arrow between two annotator boxes is labelled agreement rate. An arrow from the majority label to the model is labelled the model's score. A note says the two arrows compare different pairs, and that the majority label sits above any single annotator. Annotator A right 90% of the time Annotator B right 90% of the time agreement: 82% Majority of 5 right about 99% Model scored vs the majority score
Illustrative: the two comparisons the folk rule conflates; the numbers follow the worked example in the text.

A worked budget

The arithmetic turns into a budget in four lines. A pilot batch of 300 items, each labelled by two people, shows 82 percent agreement; inverted, that is annotators at 90 percent. The model the team hopes to ship should reach the mid-nineties, so the gold set needs to be better than that by a margin that leaves the measurement meaningful: 99 percent, say. From the chart, that is five annotators per item at 90 percent, or seven if the annotators turn out to be nearer 85. The evaluation set the team wanted, 2,000 items, therefore costs 10,000 labels rather than 2,000, and that number goes in the plan before anyone argues about it.

The alternative, 10,000 items with one label each, costs the same and is worth far less, because its ceiling is 90 percent and it cannot tell a 93 percent model from a 96 percent one. The same money buys a gold set that can, if it is spent on depth per item rather than breadth. That is the decision the inversion is for.

One caution on the formula's range. The inversion only makes sense when agreement is above 50 percent, because below that the annotators are agreeing less often than two coin flips would, and the assumption that they are each better than chance has already failed. A pilot in that range is not a budgeting problem. It is a guideline problem, and the fix is to rewrite the labelling instructions before hiring anyone.

Where the arithmetic bends

The formula assumes that annotators err independently, and they do not always. When an item is genuinely ambiguous, everyone who reads it is likelier to make the same mistake, and adding annotators helps less than the binomial says. That is the honest limit of the argument, and it points at the right response: the items where five annotators split three to two are not noise to be voted away, they are a list of the cases where the labelling guideline is unclear, and they are worth more as a document than as labels.

It also assumes annotators of roughly equal quality. Real pools have a spread, and the standard remedy, weighting each annotator by their agreement with the majority on overlapping items, is a small model in its own right and works well. Both bends make the ceiling a little lower than the clean formula gives. Neither brings it anywhere near the pairwise agreement rate, which is where the folk rule put it.

What it means for a data pipeline

SocialSure's training-data work involves labelling, and the rule I apply is short. Measure agreement on a pilot, invert it, and set the annotators per item from the accuracy the gold set needs to have, which is higher than the accuracy the model is expected to reach. Never quote pairwise agreement as a ceiling on model performance, because it is a ceiling on annotators compared with each other, and the model is not one of them. And keep the split items, because they are the cheapest guideline review you will ever get.

A model that scores above agreement is not measuring noise. It is measuring against a better judge than any single annotator, which is what a majority is for.

LabellingData QualityStatistics
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS