5 July 2026 · 5 min read

A book that looks alive can be dead

A local order book fed by WebSocket deltas keeps updating while being wrong. One dropped message desynchronises every level it touched, and a live stream proves nothing.

An order book maintained from a WebSocket stream has one property that fools everyone the first time: it keeps moving. Updates arrive, levels change, the display flickers with activity, and every one of those signs of life is compatible with the book being wrong. A single dropped message desynchronises every price level the message touched, and every later update to those levels is applied to a base that no longer matches the exchange. Shayanomaly maintains books from five venues, and the second thing I learned building it, after the first thing about fees, was that a stream being alive is not evidence of a book being correct.

Snapshot plus deltas, and the gap in between

The standard design is a snapshot followed by deltas. Fetch the full book once over REST, then apply a stream of changes: this price level is now this size, this level is gone. The design is efficient and it is fragile at exactly one point, the join between the snapshot and the stream. Binance's guide to managing a local order book spells out the procedure: open the stream first and buffer its events, then fetch the snapshot, discard any buffered event whose final update id is at or below the snapshot's, and apply the rest in order. Each event carries a first and last update id, and the rule for every subsequent event is that its first id must be exactly one more than the id your book is at. If it is greater, you missed an event, and the instruction is unambiguous: discard the book and start again.

Joining a REST snapshot to a delta stream, and detecting a gap Three lanes: the local book, the REST snapshot endpoint, and the WebSocket delta stream. The stream is opened first and its events buffered. The snapshot arrives with a last update id. Buffered events at or below that id are discarded; the first event spanning it is applied. A later event whose first id skips ahead is flagged as a gap and the book is rebuilt. Local book REST snapshot Delta stream open the stream, buffer events snapshot, lastUpdateId = 1000 drop buffered events with u at or below 1000 apply the event with U at or below 1001 and u above it next event: U must equal the book's id plus one U = 1187 when the book is at 1184: a gap discard, start again
Illustrative: the procedure described in Binance's local order book guide, with example ids.

That last rule is the whole defence, and it depends on the venue sending sequence numbers. Where it does, a gap is detectable the instant it happens. Where it does not, a gap is invisible: the next event arrives, applies cleanly to the levels it names, and the book carries on, alive and wrong.

Two venues, two integrity mechanisms

Venues differ in what they give you to detect corruption, and a system that streams several of them needs a different rule per venue rather than one integrity check applied everywhere. Binance's mechanism is sequence continuity: the first and last update ids on every event, and the rule that they must chain. Kraken's is a checksum. Its book channel sends a CRC32 with every update, calculated over the top ten price levels on each side of the book regardless of the depth subscribed, so that after applying an update you compute the same checksum over your own top ten and compare. A mismatch means your book has diverged, whether or not a message was lost, and the response is again to resubscribe.

The two mechanisms catch different things. Sequence continuity catches a missing message and nothing else; if a message arrives and is applied wrongly, the sequence is intact and the book is still wrong. A checksum catches any divergence in the top of the book, from any cause, but only in the levels it covers; a corruption deep in the book is invisible to a top-ten checksum until it rises. A venue that offers neither leaves you with one honest option, which is to fetch a fresh snapshot on a timer and accept that between snapshots the book may be wrong for as long as the timer runs.

What each integrity mechanism detects A table with three rows: sequence ids, as on Binance, detect a dropped message immediately but not a misapplied one; a checksum over the top levels, as on Kraken, detects any divergence in those levels from any cause but nothing deeper; a timed resnapshot, for venues with neither, detects nothing between snapshots and corrects everything at each one. Three ways to know the book is wrong MECHANISM CATCHES MISSES Sequence ids Binance depth streams a dropped message, at once a misapplied message Checksum, top levels Kraken book channel, CRC32 any divergence in the top ten deeper levels Timed resnapshot venues with neither everything, at each snapshot everything, between them One integrity rule cannot cover five venues; each gets the check its feed supports.
Source: mechanisms as documented by Binance and Kraken; the third row is the fallback for any feed without either.

Where the messages go missing

It is worth being concrete about how a message goes missing, because the mechanisms are mundane and none of them announces itself. A WebSocket connection drops and reconnects, and the stream resumes from now rather than from where it left off; every update during the gap is gone. A client's receive buffer fills during a burst, the library's backpressure handling discards or coalesces, and the book gets the last of a run of updates without the middle. A venue's own infrastructure fails over, and the new server's sequence numbers or snapshots are not the old server's. A message arrives out of order and is applied before its predecessor, which for a delta that says "this level is now zero" followed by one that says "this level is now 5" produces a level that should be empty and is not.

Each of those leaves the stream alive. The connection is open, messages are flowing, the display is moving. Without a per-message check, the only symptom is that the book's numbers drift from the venue's, slowly at first and then, as later updates land on corrupted levels, in ways that compound. A detector reading that book sees spreads that do not exist and misses ones that do, and it has no way to know which of its inputs is lying.

The book drift score

Detection is necessary and it is not a measurement. What I wanted, once the per-venue checks were in place, was a number that said how often each venue's book had been wrong and for how long, so that the venues could be compared and the gap-handling code could be trusted or not. The number I settled on is the book drift score: per venue, the count of price levels that differ between the locally maintained book and a periodic REST snapshot, taken at an interval and kept as a rolling metric alongside the time since the last resync.

A venue whose drift score is not zero between resyncs has a gap-handling bug, whatever its sequence checks say, because the checks passed and the book still diverged. The score also puts a duration on the damage: the time between the last event that was applied and the snapshot that revealed the drift is how many milliseconds of wrong data the detector tolerated before anyone noticed, and that number goes straight into the age gate that decides whether a quote is fresh enough to act on.

Divergence of a local book after one dropped delta, with and without detection Two lines over time after a dropped message at time zero. Without detection, the count of wrong price levels climbs steadily as later updates are applied to a corrupt base, reaching dozens within seconds. With a checksum or sequence check, the count jumps at the drop and falls to zero within one resync interval. Alive, and increasingly wrong Wrong price levels after one dropped delta; an illustrative model no detection with a check 40 30 20 10 0 drop +1 s +2 s +3 s +4 s Time after the dropped message resync, back to zero still climbing
Illustrative: the shape of divergence under a simple model in which later updates touch levels the dropped message had already changed; not a measured trace.

What the score changed

Three things, in the engine. Each venue got its own integrity rule, matched to what its feed provides, instead of one rule that was right for one venue and decorative for the rest. The resync interval per venue was set from the measured drift rather than from a guess: a venue whose score stayed at zero could be snapshotted rarely, and one whose score twitched got snapshotted often until the cause was found. And the detector stopped trusting liveness. A book is usable when its drift score is zero and its last verified update is younger than the venue's interval, and a book that is merely moving does not qualify.

The general lesson is older than order books. Any local replica fed by a stream of changes has this property: it can be arbitrarily wrong while appearing to be fully up to date, because the appearance comes from the stream and the correctness comes from the base the stream is applied to. Liveness is a fact about the connection. Correctness is a fact about the replica, and it has to be measured against the source, on a schedule, by something other than the stream that is fooling you.

WebSocketsOrder BooksCCXT
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS