5 August 2026 · 5 min read

Your backtest saw the next tick

Most leakage in time-ordered data is not a leaked column but a leaked clock. Features joined by when an event happened, not when your system could have known it, leak the future.

The first arbitrage detector I evaluated on recorded data was extraordinary. It found spreads everywhere, it closed them at the quoted prices, and its backtested returns were the kind that make you check the code for a sign error. There was no sign error. There was a clock error, and it is the same clock error that sits inside most leaky models on time-ordered data: the features were aligned by the time an event happened, not by the time my system could have known about it.

Two timestamps on every event

Every event in a time-ordered system has two times. The first is when it happened: the exchange matched the trade, the sensor read the value, the customer clicked the button. The second is when it became knowable to you: when the message arrived over the wire, when the batch job ran, when the record was written. The gap between them is latency, and it is never zero. Shayanomaly streams order books from five venues, and each message carries the exchange's timestamp and a local receipt time, and the two are never equal. The hftbacktest documentation builds its whole data model on this distinction, with an exchange timestamp and a local timestamp on every row and the feed latency defined as the difference.

One event, two timestamps, and the window in between A horizontal timeline for a single price update. A marker labelled happened-at, when the exchange matched it. A later marker labelled knowable-at, when the message arrived locally. Between them, a shaded window in which a backtest aligned by the exchange time believes it knew the price, and did not. A decision-time marker inside the window shows the leak. The window your backtest lived in A single price update as two moments, not one timestamps the leak happened-at exchange timestamp knowable-at local receipt time decision here: a price not yet received safe to use from here
Illustrative: the two-timestamp model described in the hftbacktest data documentation, drawn for one event.

A backtest that aligns everything by the exchange timestamp is a backtest in which every decision is made with information that had not yet arrived. The spreads it finds are real in the sense that the two venues really did quote those prices at those instants. They are not real in the sense that matters, which is that no system connected to those venues could have seen both quotes at the instant the backtest assumes. The detector saw the next tick, and it saw it on every trade.

The knowable-at clock

The fix generalises far beyond trading, and I keep it as a rule about tables rather than a rule about markets. Every feature carries two timestamps: happened-at and knowable-at. A training row, or a backtest decision, may use only features whose knowable-at precedes the row's decision time. Not its happened-at. Its knowable-at.

Stated that way, the familiar kinds of leakage are all the same mistake. A random train-test split on time-ordered data puts rows whose knowable-at is in the future into training. A feature built from a full-series statistic, a mean or a standard deviation over the whole history, has a knowable-at at the end of the series, and every row before that is using it early. A label that is corrected weeks after the event, a chargeback, a cancelled order, a revised earnings figure, has a knowable-at weeks after its happened-at, and a model trained on the corrected label learns from a future it will not have in production. Order books aligned by exchange time are the fast version of the same thing.

The check is three lines of SQL against any feature table that carries both columns, and I run it as a test.

SELECT count(*) AS violations
FROM training_rows r JOIN feature_values f USING (entity_id)
WHERE f.knowable_at > r.decision_time
  AND f.happened_at <= r.decision_time;

The second condition is the interesting one. Features whose happened-at is also after the decision time are the obvious future, and most pipelines already exclude them. The rows this query finds are the ones that happened before the decision and became knowable after it: the slow labels, the late-arriving messages, the statistics computed at the end. Those are the leaks nobody looks for, because by the happened-at clock they look fine.

Four kinds of leakage, one clock error A two column table. Left, the familiar name: random split, full-series feature, corrected label, exchange-time alignment. Right, the same thing stated by the knowable-at clock: future rows in training; a feature knowable only at the end of the series; a label knowable weeks after the event; a price knowable only after network latency. Different names, the same mistake THE FAMILIAR NAME BY THE KNOWABLE-AT CLOCK Random train-test split Future rows placed in training Full-series statistic as a feature Knowable only at the end of the series Corrected or revised label Knowable weeks after it happened Exchange-time alignment Knowable only after the feed latency Model with a training cutoff Knowable-at is the training corpus A chronological split alone catches only the first row.
Illustrative: the mapping this post uses; each row is a case the three-line check finds.

Why the chronological split is not enough

The standard defence against leakage on time-ordered data is to split by time: train on everything before a date, test on everything after. It is necessary and it is the first row of the table. It does nothing about the other rows. A feature computed over the whole training period, a mean of the last year's prices, say, is used by every training row including the ones at the start of the year, whose knowable-at for that feature was months in their future. The model learns a relationship between a row and a statistic the row could not have seen, the relationship holds in the test period too because the statistic was computed there as well, and the chronological split passes the leak straight through.

The same happens with labels that are revised. A fraud label finalised sixty days after the transaction is, for the last sixty days of the training window, a label from the test period. A chronological split at the boundary counts those rows as clean. The knowable-at clock counts them as leaks, and it is right, because in production the model will have to score today's transaction without knowing what the investigators will decide in two months.

The newest version of the old bug

The last row of that table is the one that has arrived with language models, and it is worth spelling out because it is the same clock error at a new scale. A model's knowable-at is its training corpus, which extends up to a cutoff date, and a backtest that asks the model to forecast events before its cutoff is asking it to recall rather than to predict. A 2026 study of temporal leakage in LLM backtesting shows that the usual check, comparing scores before and after the cutoff, is not enough: models legitimately know more about the period near their cutoff, so recency looks like leakage and leakage looks like recency, and the authors argue that only a matched clean control separates the two.

That is the knowable-at clock applied to a feature whose knowable-at is fuzzy, and the lesson is the same as for the order book. If you cannot say when the system could have known something, you cannot say whether the backtest is honest, and the default assumption should be that it is not.

What the fix looked like

For the arbitrage detector the fix was mechanical once the clock was named. Every message got its receipt time as the primary time, the exchange time was kept as a field for measuring latency, and the backtest replayed messages in receipt order, so that at every decision the detector saw exactly the set of quotes that had arrived by then. The spreads shrank, the returns fell to something believable, and the thing that remained was the actual edge: the moments when one venue's message arrived materially before another's, which is a property of network paths rather than of prices, and which is measurable.

Backtested edge under exchange-time and receipt-time alignment Two horizontal bars. Aligned by exchange time, the apparent spread captured per trade is large. Aligned by receipt time, it is a small fraction of that. The difference is the part of the edge that never existed for any real system. How much of the edge was the clock Apparent captured spread per trade, two alignments; an illustrative ratio By exchange time 100 By receipt time a small fraction The grey bar is what the detector reported. The gold bar is what any system connected to the venues could actually have captured. The gap between them was never edge. It was latency, counted as profit. drawn as a ratio; no measured returns are claimed
Illustrative: the shape of the correction, not measured results from any strategy.

Two columns, always

The rule I carry from this into every time-ordered dataset is a schema rule, and it is cheap: every feature table has both columns, happened-at and knowable-at, and the second is never defaulted to the first. For a sensor reading, knowable-at is when the reading was written to the store. For a label, it is when the label became final. For a market message, it is the local receipt time. For anything computed, it is the time the computation ran, over data that itself had knowable-at values before it.

Then the three-line check runs in CI, against every feature table, and a count above zero fails the build. It has caught more leakage than any chronological split ever did, because the chronological split protects against one row of that table and the clock protects against all of them.

EvaluationData LeakageTime Series
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS