Your backtest saw the next tick
Most leakage in time-ordered data is not a leaked column but a leaked clock. Features joined by when an event happened, not when your system could have known it, leak the future.
The first arbitrage detector I evaluated on recorded data was extraordinary. It found spreads everywhere, it closed them at the quoted prices, and its backtested returns were the kind that make you check the code for a sign error. There was no sign error. There was a clock error, and it is the same clock error that sits inside most leaky models on time-ordered data: the features were aligned by the time an event happened, not by the time my system could have known about it.
Two timestamps on every event
Every event in a time-ordered system has two times. The first is when it happened: the exchange matched the trade, the sensor read the value, the customer clicked the button. The second is when it became knowable to you: when the message arrived over the wire, when the batch job ran, when the record was written. The gap between them is latency, and it is never zero. Shayanomaly streams order books from five venues, and each message carries the exchange's timestamp and a local receipt time, and the two are never equal. The hftbacktest documentation builds its whole data model on this distinction, with an exchange timestamp and a local timestamp on every row and the feed latency defined as the difference.
A backtest that aligns everything by the exchange timestamp is a backtest in which every decision is made with information that had not yet arrived. The spreads it finds are real in the sense that the two venues really did quote those prices at those instants. They are not real in the sense that matters, which is that no system connected to those venues could have seen both quotes at the instant the backtest assumes. The detector saw the next tick, and it saw it on every trade.
The knowable-at clock
The fix generalises far beyond trading, and I keep it as a rule about tables rather than a rule about markets. Every feature carries two timestamps: happened-at and knowable-at. A training row, or a backtest decision, may use only features whose knowable-at precedes the row's decision time. Not its happened-at. Its knowable-at.
Stated that way, the familiar kinds of leakage are all the same mistake. A random train-test split on time-ordered data puts rows whose knowable-at is in the future into training. A feature built from a full-series statistic, a mean or a standard deviation over the whole history, has a knowable-at at the end of the series, and every row before that is using it early. A label that is corrected weeks after the event, a chargeback, a cancelled order, a revised earnings figure, has a knowable-at weeks after its happened-at, and a model trained on the corrected label learns from a future it will not have in production. Order books aligned by exchange time are the fast version of the same thing.
The check is three lines of SQL against any feature table that carries both columns, and I run it as a test.
SELECT count(*) AS violations
FROM training_rows r JOIN feature_values f USING (entity_id)
WHERE f.knowable_at > r.decision_time
AND f.happened_at <= r.decision_time;
The second condition is the interesting one. Features whose happened-at is also after the decision time are the obvious future, and most pipelines already exclude them. The rows this query finds are the ones that happened before the decision and became knowable after it: the slow labels, the late-arriving messages, the statistics computed at the end. Those are the leaks nobody looks for, because by the happened-at clock they look fine.
Why the chronological split is not enough
The standard defence against leakage on time-ordered data is to split by time: train on everything before a date, test on everything after. It is necessary and it is the first row of the table. It does nothing about the other rows. A feature computed over the whole training period, a mean of the last year's prices, say, is used by every training row including the ones at the start of the year, whose knowable-at for that feature was months in their future. The model learns a relationship between a row and a statistic the row could not have seen, the relationship holds in the test period too because the statistic was computed there as well, and the chronological split passes the leak straight through.
The same happens with labels that are revised. A fraud label finalised sixty days after the transaction is, for the last sixty days of the training window, a label from the test period. A chronological split at the boundary counts those rows as clean. The knowable-at clock counts them as leaks, and it is right, because in production the model will have to score today's transaction without knowing what the investigators will decide in two months.
The newest version of the old bug
The last row of that table is the one that has arrived with language models, and it is worth spelling out because it is the same clock error at a new scale. A model's knowable-at is its training corpus, which extends up to a cutoff date, and a backtest that asks the model to forecast events before its cutoff is asking it to recall rather than to predict. A 2026 study of temporal leakage in LLM backtesting shows that the usual check, comparing scores before and after the cutoff, is not enough: models legitimately know more about the period near their cutoff, so recency looks like leakage and leakage looks like recency, and the authors argue that only a matched clean control separates the two.
That is the knowable-at clock applied to a feature whose knowable-at is fuzzy, and the lesson is the same as for the order book. If you cannot say when the system could have known something, you cannot say whether the backtest is honest, and the default assumption should be that it is not.
What the fix looked like
For the arbitrage detector the fix was mechanical once the clock was named. Every message got its receipt time as the primary time, the exchange time was kept as a field for measuring latency, and the backtest replayed messages in receipt order, so that at every decision the detector saw exactly the set of quotes that had arrived by then. The spreads shrank, the returns fell to something believable, and the thing that remained was the actual edge: the moments when one venue's message arrived materially before another's, which is a property of network paths rather than of prices, and which is measurable.
Two columns, always
The rule I carry from this into every time-ordered dataset is a schema rule, and it is cheap: every feature table has both columns, happened-at and knowable-at, and the second is never defaulted to the first. For a sensor reading, knowable-at is when the reading was written to the store. For a label, it is when the label became final. For a market message, it is the local receipt time. For anything computed, it is the time the computation ran, over data that itself had knowable-at values before it.
Then the three-line check runs in CI, against every feature table, and a count above zero fails the build. It has caught more leakage than any chronological split ever did, because the chronological split protects against one row of that table and the clock protects against all of them.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS