21 August 2026 · 6 min read

A citation is a claim about a passage

A citation marker asserts that this passage supports this sentence, and it fails on its own schedule. Make the model commit to spans, check them outside the model, drop the rest.

A citation looks like evidence and is actually a second claim. The first claim is the sentence: the refund window is fourteen days. The second is the marker after it: passage 3 says so. Both can be wrong, and they fail independently, on different schedules, for different reasons. When I built grounded citation into the helpdesk assistant at CRIS, the thing I had not appreciated was that the marker needed exactly the same scrutiny as the sentence, and that giving it that scrutiny changes what the model is asked to produce.

Citations fail on their own schedule

The evidence that citation markers are unreliable is not anecdotal. The ALCE benchmark from Gao and colleagues measures citation recall, whether the cited passages entail the sentence, and citation precision, whether each cited passage is actually needed, and found the best models of its time far from full support on long-form questions. Wallat and colleagues went further in Correctness is not Faithfulness in RAG Attributions, distinguishing a citation that happens to point at a supporting passage from one that reflects how the answer was actually produced. They report that 57 percent of citations from a RAG-optimised model showed unfaithful behaviour: the model answered from its own memory and then found a passage to attach, a pattern they call post-rationalisation.

Unfaithful citations in a RAG-optimised model A stacked bar: 57 percent of citations judged unfaithful, the model answering from memory and attaching a passage afterwards, and 43 percent faithful. A note explains that a citation can point at a passage that does support the sentence and still be unfaithful. The marker can be right for the wrong reason Share of citations by faithfulness, RAG-optimised model, Wallat et al. 57% unfaithful 43% faithful answer first, passage attached after answer produced from the passage An unfaithful citation can still point at a passage that supports the sentence. It is a correct marker on a claim the passage did not produce.
Source: Wallat et al. (2024), presented at ICTIR 2025.

That distinction matters for a helpdesk more than it sounds. A post-rationalised citation is right on the day the model's memory agrees with the manual and wrong on the day the manual changes, because the model was never reading the manual. The citation looked like grounding and was decoration.

A document citation cannot be checked at scale

The usual citation is a document identifier: passage 3, or a URL. To verify it, a person opens passage 3 and reads it against the sentence. That is fine for a demonstration and impossible for a system answering thousands of questions a day. There is no cheap mechanical check for "this document supports this sentence" that does not itself involve a model, and a model checking a model's citations is the same problem with a second opinion.

A span citation is different. If the model has to quote the exact words from the passage that support the sentence, two cheap checks become possible. The first is string matching: does the quoted span actually occur in the cited passage, character for character. That check is free, deterministic and cannot be argued with. The second is entailment: does the span, read on its own, support the sentence. That check needs a model, but a small one, given a short span and a short sentence, and it is a far narrower question than "does this document support this answer".

The no-span, no-claim rule

So the rule I built is this. Every sentence in a grounded answer carries a passage identifier and a verbatim span. A checker outside the model confirms the span exists in that passage by string match, and an entailment model confirms that the span supports the sentence. A sentence that fails either check is removed from the answer before it is shown. Not flagged, not marked uncertain, removed. The answer's length becomes a function of the evidence: an answer with three supported sentences is three sentences long, and an answer with none is a decline.

The four gates a sentence passes before it is shown Five boxes left to right: generate a sentence with a passage id and a quoted span; extract the span; string-match the span against the passage; run an entailment check of the span against the sentence; show or drop. A note under the third gate says it is deterministic and free. Generate id plus span Extract the quoted span String match in that passage Entailment span supports it Show or drop deterministic and free A sentence that fails any gate is removed; the answer shrinks to its evidence.
Illustrative: the pipeline as designed; the string-match gate is the one that costs nothing and catches the most.

The string-match gate does more work than it looks. A model that has post-rationalised, answering from memory and reaching for a passage afterwards, tends to quote loosely: it paraphrases the passage, or quotes something close to it, or quotes a span from a different passage than the one it cited. Exact matching catches all of that. The model is being asked to prove it read the passage by reproducing its words, and a model that did not read it cannot.

The entailment gate catches the remainder: spans that are real but do not say what the sentence says. "Refunds are processed within fourteen days" does not support "the refund window is fourteen days", and a span-level entailment model, asked only about those two short strings, is good at seeing the difference.

What it costs

Two things, and both are worth stating plainly. The first is that answers get shorter, sometimes to nothing, and a product team has to be willing to show a two-sentence answer where the model would have written six. The trade is that the two sentences are checkable, and the four that were dropped were the ones that would have been wrong on the day the manual changed.

Answer length against the entailment threshold, a model A falling line: as the entailment threshold rises from lenient to strict, the average number of sentences surviving in an answer falls from about six to about two. A marker shows the threshold at which sentences that would fail when the source changes are removed. Shorter, and checkable Average sentences surviving per answer as the entailment gate tightens; an illustrative curve 8 6 4 2 0 lenient moderate strict Entailment threshold post-rationalised sentences fall away here 2 sentences
Illustrative: a curve drawn to show the trade the rule makes; the sentence counts are a model, not measurements.

The second cost is latency. Two extra passes, one deterministic and one a small model, run after generation and before display. The string match is microseconds. The entailment check on a handful of short pairs is tens of milliseconds on a small model, which for a helpdesk answer read by a person is invisible, and for a pipeline that must stream tokens as they are produced is a real constraint, because the gate needs the whole sentence before it can pass it. For the streamed case the answer is to gate per sentence as each completes, which delays each sentence by one check rather than the whole answer by all of them.

What the gates miss

The rule is not a guarantee of truth, and it is worth saying which failures survive it. A span can be real, support the sentence, and still be the wrong span: the passage says the refund window is fourteen days for one product and thirty for another, the model quotes the fourteen, and the entailment gate passes because the quoted words do support the sentence. The gates check that the answer came from the passage. They do not check that the passage was the right one, which is a retrieval problem and belongs to the retriever's evaluation rather than to the citation checker.

A span can also be real and the passage itself wrong, out of date, or superseded by a later document. The gates verify the chain from sentence to passage, not the passage's standing in the corpus, and a helpdesk with stale manuals will produce faithfully cited stale answers. That is a better failure than an unfaithful one, because the stale passage is named and can be fixed at the source, but it is a failure, and the corpus needs its own hygiene.

What the gates do guarantee is narrower and more valuable than it looks: every sentence shown to the reader has a named passage and a quoted span that a person can check in seconds. The failure modes that remain are the ones a person can find by following the citation. The ones removed are the ones nobody could have found, because there was nothing at the end of the marker.

What it changes about the prompt

Asking for spans changes the model's job in a way that helps before the gates even run. A model told to quote the exact supporting words for every sentence is a model that has been told, in effect, to read the passages first. It is harder to post-rationalise a quotation than a citation, because the quotation has to come from the text. The Wallat result says that a marker on its own does not make the model read; the span requirement makes reading the shortest path to a passing answer.

That is the design in one sentence: make the cheapest way to produce a citation the honest way, then check it anyway. A citation is a claim about a passage. Treat it like any other claim the model makes, and refuse to show it until it has been verified by something that is not the model.

RAGGroundingCitations
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS