The judge shares the defendant's blind spots
An LLM judge is trained on nearly the same data as the model it grades, so their errors are correlated. Judge scores inflate on the hard items; sample the calibration set there.
Using one language model to grade another is now the default way to evaluate anything that does not have a single right answer, and it works well enough on average that the average has become the whole story. It should not be. A judge model is trained on nearly the same data as the model it grades, by nearly the same methods, toward nearly the same objective, and so its errors are not independent of the generator's. Where the generator is confidently wrong, the judge is likely to be wrong in the same way, for the same reason, and to give the wrong answer a high score. The measured accuracy is therefore inflated exactly on the items that matter, the hard ones, and the small human-labelled sample that is supposed to calibrate the judge will miss the inflation unless it is drawn deliberately from where the inflation lives.
The average is fine
The evidence for judges is real. Zheng and colleagues, in the MT-Bench and Chatbot Arena paper, found that a strong judge agreed with human preferences over 80 percent of the time, a level that matched how often humans agreed with each other, and they named the biases they found on the way: a preference for the first of two answers, a preference for longer answers, a preference for the judge's own style, and limited ability to grade reasoning. Every subsequent judge-based evaluation leans on that 80 percent, and it is an honest number. It is also an average over items of every difficulty, and averages are where correlated errors hide.
The grid shows why. A judge's error can go two ways. It can mark a right answer wrong, which lowers the score, gets noticed, and gets investigated. Or it can mark a wrong answer right, which raises the score and is invisible, because the score is exactly the count of the left column and nothing in it distinguishes the top cell from the bottom. Independent errors would fill the bottom-left cell at random, in proportion to the judge's overall error rate, and a random human sample would catch them. Correlated errors fill it systematically, on the items where both models share a misconception, and a random sample rarely lands there because those items are a small fraction of the whole.
Why the errors are correlated
The correlation has a documented mechanism. Panickssery, Bowman and Feng showed in LLM Evaluators Recognize and Favor Their Own Generations that a judge scores its own outputs higher than others' while humans rate them equal, and that the strength of this self-preference is linearly related to how well the judge can recognise its own output, a relationship they established by manipulating the recognition ability directly. Self-preference is the sharpest case of a general one. A judge and a generator that share pretraining data share the errors in that data; ones that share a fine-tuning recipe share the stylistic tells that recipe produces; ones that share an architecture share the reasoning failures the architecture has. The judge does not need to be the same model as the generator for the errors to correlate. It needs to have learned the world the same way.
I call the failure mode correlated blindness: judge and generator fail on the same items for the same reason. Its signature is a confident, fluent, wrong answer that the judge reads as confident, fluent and right, and the reason it is dangerous is that confidence and fluency are precisely the surface features a judge uses when it cannot verify the substance. On an easy item the judge can verify, and correlation does not matter. On a hard item it cannot, and it falls back to the features it shares with the generator.
The agreement-weighted sample
The standard defence is a human-labelled calibration set: label a few hundred items by hand, compare the judge's verdicts to the labels, and report the judge's accuracy as the confidence in its scores. That defence assumes the calibration set samples the failure mode, and a random sample does not. The correlated errors live in one cell, wrong answer with a high judge score, and within that cell they concentrate on the items where the generator was confident, because confident wrong answers are the ones that share the judge's blind spot. A random sample of a few hundred items from a set where most answers are right lands in that region a handful of times.
So the calibration sample should be weighted toward it. I call this the agreement-weighted sample: human-label a set drawn preferentially from items where the judge scored high and the generator was confident, since that is the cell where correlated errors hide and the cell random sampling rarely reaches. The weighting is not subtle; it can be as simple as taking every item where the judge gave a top score and the generator's own confidence was in its top quartile, and labelling as many of those as the budget allows, with a smaller random slice for comparison. The judge's accuracy on that weighted set is the number that says whether its scores on hard items can be believed, and it is usually lower than the accuracy on the random set, which is the point.
Two practical notes on building the weighted set. The generator's confidence is not always exposed as a number, and when it is not, a proxy works: the absence of hedging language in the answer, or agreement between several samples of the same answer, both of which track the confident-wrong failure well enough to weight by. And the judge's own confidence can be used the same way in reverse, by preferring items where the judge's score was high but its stated reasoning was thin, which is a cheap signal that it graded on surface features. Neither proxy needs to be exact. The sample only has to land in the corner more often than a random one does, and a random one lands there almost never.
Applying it to a helpdesk evaluation
The place I reason about this concretely is a helpdesk assistant's evaluation set, where a judge grades answers for correctness and grounding against retrieved documents. The generator's confident wrong answers are the ones that read like the documentation, cite a plausible section and state a procedure that does not exist. The judge, which learned what documentation sounds like from the same corpus, is exactly as fooled by that as the generator was, and it scores the answer high. A random calibration sample of two hundred items, from a set where the assistant is right most of the time, finds two or three of those. The weighted sample, drawn from items with a top judge score and a confident generator, finds them at a rate that says something, and the honest evaluation reports both numbers: the judge's accuracy overall, and its accuracy on the weighted set, with the second labelled as the confidence you should have in the scores on hard questions.
The rule generalises. Any time one model grades another, ask what they share, assume their errors are correlated in proportion to it, and draw the human sample from the cell where correlated errors accumulate. The 80 percent is true. It is just not the number that describes the answers you are most worried about.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS