Your test set has been seen before
Near-duplicates across the split boundary are the commonest reason a held-out score is not a generalisation score. Dedupe before you split, and report the overlap next to accuracy.
A held-out test set is supposed to measure how a model does on data it has not seen. It measures that only if the test data really is unseen, and in most real datasets a surprising share of it is not: not identical to a training example, but close enough that the model has, in every sense that matters, seen it before. The split was random, the duplicates were in the source, and the random split preserved them on both sides of the line. The score that comes out is not a generalisation score. It is partly a memorisation score, and it is higher than the truth.
How much of a test set is a sibling
The numbers on well-known datasets are not small. Barz and Denzler's Do We Train on Test Data? searched CIFAR for near-duplicates and found that 3.3 percent of CIFAR-10's test images and 10 percent of CIFAR-100's have a duplicate in the training set; when they replaced those images with fresh ones from the same domain, classification accuracy fell by between 9 and 14 percent relative to the reported figures. Lee and colleagues' Deduplicating Training Data Makes Language Models Better found the same structure in text: over 4 percent of the validation sets of standard language-modelling corpora had an approximate duplicate in the training data, including one 61-word sentence that appeared over sixty thousand times in C4.
Those are curated academic datasets, assembled by people who cared. A dataset scraped or collected inside a company is worse, and for reasons that are easy to name: the same page fetched twice on different days, an image and its thumbnail, a record and its retried submission, a document and the copy someone pasted into a ticket, an augmentation that was applied before the split instead of after. Every one of those puts a sibling on each side of the boundary, and a random split does not know.
Dedupe before you split, and look for siblings
The fix has two halves, and the order is the important part. Deduplication has to happen before the split, because a split of a duplicated source is a duplicated split. And the deduplication has to look for near-duplicates, because exact hashing catches almost none of the cases above: the thumbnail is not byte-identical to the image, the retried record has a new timestamp, the pasted copy has different whitespace.
Near-duplicate detection is a solved problem at the scale most teams need. For text, minhash over shingles or embedding similarity with a threshold; for images, a perceptual hash or an embedding from a pretrained encoder; for records, a similarity over the fields that matter with the timestamps and identifiers stripped out. The threshold has to be chosen, and the honest way to choose it is to sample pairs at several thresholds and look at them: the level at which a person says "these are the same thing" is the level, and it is different for every dataset.
Why the score moves so much
A few percent of siblings sounds like it should move the score by a few percent, and the CIFAR result says it moves it by more. The reason is that siblings are not a random sample of the test set. They are the items the model finds easiest, because it has effectively seen them, and they are concentrated in exactly the classes and regions where the model is otherwise weakest, because those are where the source had the most redundancy: rare classes padded with near-copies, hard cases collected twice. Remove them and the test set gets harder in the places that were propping up the number.
The other reason is what the siblings were doing during training. A near-duplicate of a test item in the training set is not only a free point at test time; it is a training example the model was rewarded for memorising, and a model that has learnt to memorise redundant items generalises a little worse on everything else. Lee and colleagues' result, that deduplicated training data produces models that memorise less and reach the same accuracy in fewer steps, is the training-side half of the same finding. The sibling in the test set inflates the score; the sibling in the training set depresses the model. Removing both is what the pipeline order buys.
The sibling scan
Deduplication before the split is a pipeline habit, and habits lapse. What keeps the habit honest is a measurement that runs every time and reports next to accuracy, and the one I use is what I call the sibling scan. Before any metric is read, embed or hash every item in the test set and every item in the training set, find each test item's nearest training neighbour, and report the fraction of test items whose nearest neighbour is above the similarity cutoff. That fraction is the sibling rate, and it sits on the evaluation report as a first-class number beside the score it qualifies.
The report line is the point. An accuracy with a sibling rate beside it is a score that can be read: 94 percent with 6 percent siblings is a model whose true held-out accuracy is somewhat lower than 94 and whose evaluation set needs cleaning. An accuracy with no sibling rate is a number whose meaning depends on a pipeline habit nobody can see. The scan costs a nearest-neighbour search over the test set, which for any evaluation set a team actually uses is seconds, and it has caught leaks in pipelines that had been deduplicating carefully, because the deduplication had been happening after the split.
Where the cutoff comes from
The cutoff has two honest sources, and I use both. The first is a small hand-labelled sample: draw pairs across a range of similarities, have a person say which are the same thing, and put the cutoff where the answers flip. The second is the random-pair baseline of the embedding model itself, the similarity that unrelated items have under this model on this data, which I wrote about in the post on similarity thresholds: a cutoff well above that baseline is a cutoff that means something, and a cutoff copied from another dataset is a cutoff that means whatever that dataset's baseline meant.
SocialSure's training-data work is where this habit lives for me, and the reason it lives there is simple. A pipeline that assembles data from scrapes, from customer uploads and from augmentation produces siblings at every stage, and the first time a model scored suspiciously well on a held-out set was the last time the held-out set was trusted without a sibling rate next to it. The score is not the measurement. The score and the overlap together are.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS