Two models your eval cannot tell apart
A benchmark's item count sets its resolution. On HumanEval a six-point gap between two models is a tie, and most leaderboard gaps are smaller than that.
Two checkpoints, one evaluation set of two hundred questions. Checkpoint A scores 81, checkpoint B scores 78. The team ships A. Nobody asks whether 81 and 78 are different numbers, because they look different, and looking different is enough at four in the afternoon.
They are not different. At two hundred items, a three-point gap is well inside the noise of the measurement, and the decision to ship A was a coin toss with a spreadsheet attached. This post is about the size of that noise, which turns out to be a fixed property of the benchmark that you can look up before you run anything.
A benchmark's resolution
An accuracy score is a proportion: correct items over total items. If the benchmark's questions are a sample from some larger population of questions you care about, and Evan Miller's Adding Error Bars to Evals argues that this is the only sensible way to read them, then the score has a standard error like any other proportion: the square root of p times one minus p, divided by the item count.
At an accuracy of 80% the numerator is fixed, so the standard error depends only on the number of items, and the 95% interval is about two standard errors either side. Work that through for the benchmarks people actually quote. HumanEval has 164 problems, from the paper that introduced it. GSM8K's test split has 1,319, per Cobbe et al.. SWE-bench Verified has 500, per OpenAI's announcement. ARC-Challenge's test set has 1,172, per Clark et al.. MMLU's test set has 14,042, per Hendrycks et al..
Read the first bar. A model that scores 80 on HumanEval could plausibly be a 74 model or an 86 model, on the same task, with a different draw of 164 problems. That is not a criticism of HumanEval. It is what 164 items can resolve, and it has been true since the day the benchmark was published.
Two scores, one gap
Comparing two models is worse than scoring one, because both scores carry noise. If the two models were evaluated on different draws of questions, or if you only kept the two totals, the standard error of the difference is the single-score error times the square root of two. Call the 95% half-width of that difference the tie zone: any gap smaller than it cannot be ranked. On HumanEval the tie zone at 80% accuracy is about 8.7 points. On SWE-bench Verified it is 5.0. On GSM8K it is 3.1. On MMLU it is 0.9. A great many of the gaps announced with adjectives on leaderboards live inside those bands.
There is a way to shrink the zone, and it costs nothing but bookkeeping. Evaluate both models on the same items and keep the per-item outcomes. Then the comparison is paired: for each question you know whether A and B both passed, both failed, or split. Only the split items carry information about the difference, and when the two models agree on most items, which similar checkpoints do, the paired standard error can be a fraction of the unpaired one. Miller's paper works through the paired formula and recommends it as the default; it is the difference between a coin toss and a measurement.
The tie zone as a rule
Two habits follow. The first is for reading other people's numbers: report a gap smaller than the benchmark's tie zone as a tie, whatever the leaderboard's sort order says. A leaderboard sorts on the point estimate because a table needs an order, not because the order is known. A model that is "second" by half a point on MMLU may be first; the table cannot tell, and neither can you.
The second habit is for your own evaluations: decide the gap you need to resolve before you build the set, and size the set to it. To tell a three-point difference from zero at 95% confidence, unpaired, at around 80% accuracy, you need roughly 1,400 items. To resolve one point you need about 12,000. Pairing brings both numbers down, sometimes by a lot, but only if the outcomes are kept per item, which is a decision made on the day the evaluation code is written and expensive to reverse later.
The noise the item count does not cover
The item count sets the floor on the noise, and two other sources sit on top of it. The first is the model's own randomness. A sampled answer at any temperature above zero is one draw, so a score of 80 on one run is itself a sample of the model's behaviour on those questions, and rerunning the same checkpoint on the same items can move the total by a point or two before any question changes. Miller's recommendation is to score several samples per question and average them, which is also the only way the paired comparison stays honest, since a single lucky draw on a split item would otherwise decide the gap.
The second is structure in the items. Benchmarks built from passages, with several questions per passage, or from problem families with shared templates, do not contain as many independent items as they have rows. Questions that share a passage tend to be answered right or wrong together, and the standard error must be computed on the clusters rather than the rows, which widens the interval further. A reading-comprehension set with 2,000 questions over 400 passages has the resolution of something closer to 400 items than 2,000, and a leaderboard that reports it as 2,000 is quietly overstating its own precision.
Neither correction changes the direction of the argument. They both say the tie zone is wider than the first chart shows, never narrower.
What I do with a small set
When I was evaluating a helpdesk assistant at CRIS, the internal question set was the size internal sets always are: small, hand-written, and precious. The temptation with a set like that is to treat every point of movement as a signal, and the tie zone says that most of the movement is weather.
So the rules I use now are short. Keep the per-item outcomes for every run, so that any two checkpoints can be compared paired rather than by their totals. Score more than one sample per question when the decoding is not deterministic. Decide the smallest gap that would change a decision, and compute the item count that gap needs; if the set is too small to resolve it, say so in the report rather than rounding the noise into a verdict.
And when two checkpoints tie, which on a small set is most of the time, choose on the things the eval does not measure: latency, cost, and the failure modes a person can read in the transcripts. A tie is not a failure of the evaluation. It is the evaluation telling you, correctly, that the decision belongs to a different set of numbers.
A three-point gap on two hundred items is not a reason to ship A. It is a reason to write down how many items it would take to know, and to decide whether knowing is worth that many questions.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS