The last passage that mattered
Retrieval depth is set by feel and never touched, yet every extra passage is prefill and latency on every turn. Measure the rank of the last passage an answer used, then set k.
Somewhere in every retrieval pipeline is a number that was chosen on the first day and never revisited: how many passages to retrieve. Five, because the tutorial said five. Ten, because ten felt safer. The number sets the prompt length on every turn, which sets the prefill cost and the time to first token, and it also sets how much noise sits next to the evidence. When I tuned the helpdesk assistant at CRIS against a latency target, retrieval depth was one of three knobs I had, and it was the one nobody had a principled way to set. This is the way I settled on.
More is not better past a point
The intuition that more passages can only help is wrong, and the evidence for that has accumulated. Liu and colleagues' Lost in the Middle showed that models use evidence unevenly across a long context, with accuracy dropping when the relevant passage sits in the middle of the retrieved set. Cuconasu and colleagues' The Power of Noise found that a moderate share of irrelevant passages in the retrieved set substantially degrades factual accuracy, and argued that precision and reranking matter more than raw recall. A 2026 systems-level analysis of retrieval-augmented generation reports the pattern directly in one of its configurations: answer quality peaked at a retrieval depth of three and fell at higher depths, even as more of the gold evidence was retrieved.
So the curve of answer quality against k rises, peaks early, and then flattens or falls, and the passages past the peak are paid for twice: once in latency and once in noise.
That tells you the shape. It does not tell you where the knee is for your corpus, your chunk size and your questions, and the published peaks move around from study to study because those things differ. The knee has to be measured locally, and the usual way to measure it, sweeping k and re-running the whole evaluation at each value, is expensive enough that nobody does it more than once.
The last useful passage
The measurement I use instead is cheaper and more direct. For each question in the evaluation set, retrieve generously, say twenty passages, generate the answer, and then find the highest rank among the passages the answer actually used: the ones it cites, or, where citation is not enforced, the ones that contain the gold span. That rank is the question's last useful passage, and I write it as LUP. A question whose answer drew on passages ranked 1, 2 and 7 has an LUP of 7. A question whose answer used only the top passage has an LUP of 1.
The distribution of LUP across the evaluation set is the whole picture of how deep retrieval needs to go, and one number from it sets k: the 95th percentile. Retrieve that many and nineteen questions in twenty have every passage they needed. Retrieve more and you are paying prefill on every turn for the twentieth question, which the noise past the knee is likely to hurt anyway.
What the histogram tells you that k does not
The histogram is worth more than the single number, because its shape diagnoses the retriever. A tall bar at rank one and a short tail says the retriever is putting the evidence first and a small k is safe. A flat spread across ranks says the retriever cannot distinguish the relevant passage from its neighbours, and the fix is reranking or better embeddings, not a larger k. A second bump at high ranks says some class of question needs evidence the retriever ranks poorly, and it is worth looking at those questions by hand, because they are usually a chunking problem: the answer spans two chunks, or lives in a table the chunker split.
The measurement also separates two things that a sweep over k conflates. Retrieval quality is where the evidence lands in the ranking. Generation quality is what the model does with it. LUP measures the first directly, on the answers the model actually produced, and leaves the second to the usual accuracy metrics. When accuracy is poor and LUP is low, the retriever is fine and the prompt or the model is the problem. When LUP is high, no prompt will save a pipeline that is hiding its evidence at rank nine.
Two things the measurement needs to be honest
The first is that the answer's use of a passage has to be observable, and the cleanest way to make it observable is to require citations in the prompt: every sentence the model writes names the passage it came from. That is a design choice with benefits beyond this measurement, since it is also how a helpdesk answer becomes checkable by the person reading it, but it changes the model's behaviour, and a pipeline measured with citations on should run with citations on. Where citations cannot be enforced, the fallback is the gold span: the passage that contains the answer the evaluation set says is correct, whether or not the model used it. That measures where the evidence was rather than what the model touched, which is a slightly different quantity and still a good one.
The second is that the generous retrieval used for the measurement must not be the retrieval used in production, or the measurement is circular. Retrieve twenty to find out that four would have done, then run with four. If the histogram's tail grows when you check again in a month, the corpus or the questions have changed, and the number moves with them. That is the point of making it a measurement rather than a setting.
What the extra passages cost
Every passage above the 95th percentile is prefill on every turn, and prefill is the part of latency the user feels most, because it happens before the first token. A passage of three hundred tokens, at k of ten instead of four, is eighteen hundred extra tokens of prompt on every request. That is not only time to first token; on a hosted API it is billed input on every turn of every conversation, and on a self-hosted model it is accelerator time that could have served another user.
How to run it
The measurement needs three things most teams already have: an evaluation set with gold spans or a citation-enforcing prompt, a retriever that can return more than it normally does, and a way to map an answer back to the passages it used. Retrieve twenty, generate, record LUP per question, and plot. Set k at the 95th percentile, or the 90th if latency is tight and the tail is thin. Then re-run the measurement whenever the corpus, the chunker or the embedding model changes, because each of those moves the histogram and none of them announces it.
One more habit follows from having the histogram: keep it per question type. Short factual questions and long procedural ones usually have different tails, and a single k chosen for both is either too deep for the first or too shallow for the second. Where the pipeline can classify the question cheaply, it can choose k per class, and the two histograms tell it what to choose.
The number that came out of the tutorial was never the problem. The problem was that it stayed a number instead of becoming a measurement, and the measurement costs one afternoon.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS