20 April 2025 · 5 min read

The judge shares the defendant's blind spots

An LLM judge is trained on nearly the same data as the model it grades, so their errors are correlated. Judge scores inflate on the hard items; sample the calibration set there.

Using one language model to grade another is now the default way to evaluate anything that does not have a single right answer, and it works well enough on average that the average has become the whole story. It should not be. A judge model is trained on nearly the same data as the model it grades, by nearly the same methods, toward nearly the same objective, and so its errors are not independent of the generator's. Where the generator is confidently wrong, the judge is likely to be wrong in the same way, for the same reason, and to give the wrong answer a high score. The measured accuracy is therefore inflated exactly on the items that matter, the hard ones, and the small human-labelled sample that is supposed to calibrate the judge will miss the inflation unless it is drawn deliberately from where the inflation lives.

The average is fine

The evidence for judges is real. Zheng and colleagues, in the MT-Bench and Chatbot Arena paper, found that a strong judge agreed with human preferences over 80 percent of the time, a level that matched how often humans agreed with each other, and they named the biases they found on the way: a preference for the first of two answers, a preference for longer answers, a preference for the judge's own style, and limited ability to grade reasoning. Every subsequent judge-based evaluation leans on that 80 percent, and it is an honest number. It is also an average over items of every difficulty, and averages are where correlated errors hide.

Generator right or wrong against judge right or wrong, with the dangerous cell marked A two by two grid. Rows: the generator's answer is right or wrong. Columns: the judge's verdict is right or wrong. Top left, generator right and judge right: counted correctly. Bottom right, generator wrong and judge says wrong: counted correctly. Top right, generator right and judge says wrong: a visible miss that lowers the score. Bottom left, generator wrong and judge says right: the dangerous cell, an invisible inflation, and the one that correlated errors fill. The cell you cannot see from the score JUDGE SAYS RIGHT JUDGE SAYS WRONG GENERATOR RIGHT GENERATOR WRONG counted correctly right answer, right verdict a visible miss lowers the score; someone will look invisible inflation wrong answer, high score where correlated errors go counted correctly wrong answer, caught If both models fail on the same items, the highlighted cell fills and the score hides it.
Illustrative: the four outcomes of judging one answer; the score is computed from the left column and cannot tell its two cells apart.

The grid shows why. A judge's error can go two ways. It can mark a right answer wrong, which lowers the score, gets noticed, and gets investigated. Or it can mark a wrong answer right, which raises the score and is invisible, because the score is exactly the count of the left column and nothing in it distinguishes the top cell from the bottom. Independent errors would fill the bottom-left cell at random, in proportion to the judge's overall error rate, and a random human sample would catch them. Correlated errors fill it systematically, on the items where both models share a misconception, and a random sample rarely lands there because those items are a small fraction of the whole.

Why the errors are correlated

The correlation has a documented mechanism. Panickssery, Bowman and Feng showed in LLM Evaluators Recognize and Favor Their Own Generations that a judge scores its own outputs higher than others' while humans rate them equal, and that the strength of this self-preference is linearly related to how well the judge can recognise its own output, a relationship they established by manipulating the recognition ability directly. Self-preference is the sharpest case of a general one. A judge and a generator that share pretraining data share the errors in that data; ones that share a fine-tuning recipe share the stylistic tells that recipe produces; ones that share an architecture share the reasoning failures the architecture has. The judge does not need to be the same model as the generator for the errors to correlate. It needs to have learned the world the same way.

I call the failure mode correlated blindness: judge and generator fail on the same items for the same reason. Its signature is a confident, fluent, wrong answer that the judge reads as confident, fluent and right, and the reason it is dangerous is that confidence and fluency are precisely the surface features a judge uses when it cannot verify the substance. On an easy item the judge can verify, and correlation does not matter. On a hard item it cannot, and it falls back to the features it shares with the generator.

Measured accuracy against true accuracy as the error correlation between judge and generator rises, from a stated toy model Lines showing measured accuracy against true accuracy for three levels of error correlation. With independent errors, measured accuracy tracks true accuracy closely. With moderate correlation, measured accuracy sits above the diagonal, most of all in the middle of the range. With high correlation, measured accuracy stays high even as true accuracy falls, so a generator that is right half the time can score near eighty percent. The more the two models share, the less the score can fall Judge-measured against true accuracy, three error correlations 100% 50% 0 0 50% 100% True accuracy of the generator independent moderate high half right, scored near 80 grey diagonal: a perfect judge
Illustrative: a toy model in which the judge has a fixed error rate and a stated fraction of its errors fall on the generator's wrong answers; the curves are constructed to show the shape and are not measurements.

The agreement-weighted sample

The standard defence is a human-labelled calibration set: label a few hundred items by hand, compare the judge's verdicts to the labels, and report the judge's accuracy as the confidence in its scores. That defence assumes the calibration set samples the failure mode, and a random sample does not. The correlated errors live in one cell, wrong answer with a high judge score, and within that cell they concentrate on the items where the generator was confident, because confident wrong answers are the ones that share the judge's blind spot. A random sample of a few hundred items from a set where most answers are right lands in that region a handful of times.

So the calibration sample should be weighted toward it. I call this the agreement-weighted sample: human-label a set drawn preferentially from items where the judge scored high and the generator was confident, since that is the cell where correlated errors hide and the cell random sampling rarely reaches. The weighting is not subtle; it can be as simple as taking every item where the judge gave a top score and the generator's own confidence was in its top quartile, and labelling as many of those as the budget allows, with a smaller random slice for comparison. The judge's accuracy on that weighted set is the number that says whether its scores on hard items can be believed, and it is usually lower than the accuracy on the random set, which is the point.

Where a random calibration sample lands against where the agreement-weighted sample is drawn Two panels over the same population of graded items, drawn as a field with a small dense region in one corner representing confident, high-scored, wrong answers. In the left panel, random sample points are scattered evenly and almost none fall in the dense region. In the right panel, the sample is concentrated in the region where the judge scored high and the generator was confident, so the dense region is covered. Sample where the errors hide, not where the items are RANDOM SAMPLE confident, scored high, wrong AGREEMENT-WEIGHTED SAMPLE The corner is where judge and generator fail together; only the second sample reaches it.
Illustrative: the two sampling schemes over one population of graded items; the corner is the cell of the first figure that correlated errors fill.

Two practical notes on building the weighted set. The generator's confidence is not always exposed as a number, and when it is not, a proxy works: the absence of hedging language in the answer, or agreement between several samples of the same answer, both of which track the confident-wrong failure well enough to weight by. And the judge's own confidence can be used the same way in reverse, by preferring items where the judge's score was high but its stated reasoning was thin, which is a cheap signal that it graded on surface features. Neither proxy needs to be exact. The sample only has to land in the corner more often than a random one does, and a random one lands there almost never.

Applying it to a helpdesk evaluation

The place I reason about this concretely is a helpdesk assistant's evaluation set, where a judge grades answers for correctness and grounding against retrieved documents. The generator's confident wrong answers are the ones that read like the documentation, cite a plausible section and state a procedure that does not exist. The judge, which learned what documentation sounds like from the same corpus, is exactly as fooled by that as the generator was, and it scores the answer high. A random calibration sample of two hundred items, from a set where the assistant is right most of the time, finds two or three of those. The weighted sample, drawn from items with a top judge score and a confident generator, finds them at a rate that says something, and the honest evaluation reports both numbers: the judge's accuracy overall, and its accuracy on the weighted set, with the second labelled as the confidence you should have in the scores on hard questions.

The rule generalises. Any time one model grades another, ask what they share, assume their errors are correlated in proportion to it, and draw the human sample from the cell where correlated errors accumulate. The 80 percent is true. It is just not the number that describes the answers you are most worried about.

EvaluationLLM-as-Judge
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS