19 June 2026 · 5 min read

Two models your eval cannot tell apart

A benchmark's item count sets its resolution. On HumanEval a six-point gap between two models is a tie, and most leaderboard gaps are smaller than that.

Two checkpoints, one evaluation set of two hundred questions. Checkpoint A scores 81, checkpoint B scores 78. The team ships A. Nobody asks whether 81 and 78 are different numbers, because they look different, and looking different is enough at four in the afternoon.

They are not different. At two hundred items, a three-point gap is well inside the noise of the measurement, and the decision to ship A was a coin toss with a spreadsheet attached. This post is about the size of that noise, which turns out to be a fixed property of the benchmark that you can look up before you run anything.

A benchmark's resolution

An accuracy score is a proportion: correct items over total items. If the benchmark's questions are a sample from some larger population of questions you care about, and Evan Miller's Adding Error Bars to Evals argues that this is the only sensible way to read them, then the score has a standard error like any other proportion: the square root of p times one minus p, divided by the item count.

At an accuracy of 80% the numerator is fixed, so the standard error depends only on the number of items, and the 95% interval is about two standard errors either side. Work that through for the benchmarks people actually quote. HumanEval has 164 problems, from the paper that introduced it. GSM8K's test split has 1,319, per Cobbe et al.. SWE-bench Verified has 500, per OpenAI's announcement. ARC-Challenge's test set has 1,172, per Clark et al.. MMLU's test set has 14,042, per Hendrycks et al..

How far a single score can be from the truth, by benchmark Horizontal bars of the 95% half-width in accuracy points for one model scoring 80%: HumanEval with 164 items, 6.1; SWE-bench Verified with 500 items, 3.5; ARC-Challenge with 1,172 items, 2.3; GSM8K with 1,319 items, 2.2; MMLU with 14,042 items, 0.7. HumanEval is highlighted. 95% half-width of one score at 80% accuracy Accuracy points, from the binomial standard error HumanEval (164) 6.1 SWE-bench Verified (500) 3.5 ARC-Challenge (1,172) 2.3 GSM8K (1,319) 2.2 MMLU (14,042) 0.7
Source: item counts from the HumanEval, SWE-bench Verified, ARC, GSM8K and MMLU papers; half-widths computed as 1.96 times the binomial standard error at p = 0.8.

Read the first bar. A model that scores 80 on HumanEval could plausibly be a 74 model or an 86 model, on the same task, with a different draw of 164 problems. That is not a criticism of HumanEval. It is what 164 items can resolve, and it has been true since the day the benchmark was published.

Two scores, one gap

Comparing two models is worse than scoring one, because both scores carry noise. If the two models were evaluated on different draws of questions, or if you only kept the two totals, the standard error of the difference is the single-score error times the square root of two. Call the 95% half-width of that difference the tie zone: any gap smaller than it cannot be ranked. On HumanEval the tie zone at 80% accuracy is about 8.7 points. On SWE-bench Verified it is 5.0. On GSM8K it is 3.1. On MMLU it is 0.9. A great many of the gaps announced with adjectives on leaderboards live inside those bands.

There is a way to shrink the zone, and it costs nothing but bookkeeping. Evaluate both models on the same items and keep the per-item outcomes. Then the comparison is paired: for each question you know whether A and B both passed, both failed, or split. Only the split items carry information about the difference, and when the two models agree on most items, which similar checkpoints do, the paired standard error can be a fraction of the unpaired one. Miller's paper works through the paired formula and recommends it as the default; it is the difference between a coin toss and a measurement.

Paired and unpaired comparison of two models Two rows of eight items for model A and model B. In the unpaired view only the two totals, six of eight and five of eight, are kept. In the paired view each column is compared: six columns agree and two split, and only the two split columns carry information about the difference. Keep the per-item outcomes The same eight questions, read two ways Model A Model B 6 / 8 5 / 8 Unpaired: two totals, one point apart, noise from all eight items on both sides Paired: six columns agree and cancel; the difference lives in two split columns pass, both fail pass where the other failed
Illustrative: a toy set of eight items, drawn to show which columns carry information about a difference.

The tie zone as a rule

Two habits follow. The first is for reading other people's numbers: report a gap smaller than the benchmark's tie zone as a tie, whatever the leaderboard's sort order says. A leaderboard sorts on the point estimate because a table needs an order, not because the order is known. A model that is "second" by half a point on MMLU may be first; the table cannot tell, and neither can you.

The second habit is for your own evaluations: decide the gap you need to resolve before you build the set, and size the set to it. To tell a three-point difference from zero at 95% confidence, unpaired, at around 80% accuracy, you need roughly 1,400 items. To resolve one point you need about 12,000. Pairing brings both numbers down, sometimes by a lot, but only if the outcomes are kept per item, which is a decision made on the day the evaluation code is written and expensive to reverse later.

Tie zone against evaluation set size A single line falling as the item count rises: 11.1 points at 100 items, 7.8 at 200, 5.0 at 500, 3.5 at 1,000, 2.5 at 2,000, 1.6 at 5,000 and 1.1 at 10,000, for an unpaired comparison at 80% accuracy. The gap you can resolve, by item count 95% tie zone in accuracy points, unpaired, both models near 80% 12 9 6 3 0 100 200 500 1,000 2,000 5,000 10,000 Items in the evaluation set (spacing is categorical, not linear) 11.1 1.1
Illustrative: computed as 1.96 times the square root of 2 times p(1 minus p) over n, with p = 0.8; the construction follows the unpaired difference in Miller (2024).

The noise the item count does not cover

The item count sets the floor on the noise, and two other sources sit on top of it. The first is the model's own randomness. A sampled answer at any temperature above zero is one draw, so a score of 80 on one run is itself a sample of the model's behaviour on those questions, and rerunning the same checkpoint on the same items can move the total by a point or two before any question changes. Miller's recommendation is to score several samples per question and average them, which is also the only way the paired comparison stays honest, since a single lucky draw on a split item would otherwise decide the gap.

The second is structure in the items. Benchmarks built from passages, with several questions per passage, or from problem families with shared templates, do not contain as many independent items as they have rows. Questions that share a passage tend to be answered right or wrong together, and the standard error must be computed on the clusters rather than the rows, which widens the interval further. A reading-comprehension set with 2,000 questions over 400 passages has the resolution of something closer to 400 items than 2,000, and a leaderboard that reports it as 2,000 is quietly overstating its own precision.

Neither correction changes the direction of the argument. They both say the tie zone is wider than the first chart shows, never narrower.

What I do with a small set

When I was evaluating a helpdesk assistant at CRIS, the internal question set was the size internal sets always are: small, hand-written, and precious. The temptation with a set like that is to treat every point of movement as a signal, and the tie zone says that most of the movement is weather.

So the rules I use now are short. Keep the per-item outcomes for every run, so that any two checkpoints can be compared paired rather than by their totals. Score more than one sample per question when the decoding is not deterministic. Decide the smallest gap that would change a decision, and compute the item count that gap needs; if the set is too small to resolve it, say so in the report rather than rounding the noise into a verdict.

And when two checkpoints tie, which on a small set is most of the time, choose on the things the eval does not measure: latency, cost, and the failure modes a person can read in the transcripts. A tie is not a failure of the evaluation. It is the evaluation telling you, correctly, that the decision belongs to a different set of numbers.

A three-point gap on two hundred items is not a reason to ship A. It is a reason to write down how many items it would take to know, and to decide whether knowing is worth that many questions.

EvaluationStatisticsBenchmarks
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS