8 September 2026 · 5 min read

Five order books, one frame budget

A terminal on five venues receives more order-book updates than it can paint; rendering per message freezes the tab when a trader needs it. The unit of work is the frame.

A browser has one budget that matters for anything that moves: roughly sixteen milliseconds per frame at sixty frames a second, shared between everything the page wants to do. An order-book feed from a busy exchange does not know about that budget. It sends an update whenever the exchange's own cadence says to, and when a trading terminal is subscribed to five venues at once, the updates arrive faster than frames do, at exactly the moments, a spike or a cascade of liquidations, when the trader most needs the screen to keep up. Rendering per message is the natural first implementation and it guarantees a frozen tab in those moments. The fix is to change the unit of work from the message to the frame.

How fast the feeds actually are

The venues publish their cadences, and they differ by an order of magnitude. Binance's diff depth stream offers updates every 1,000 milliseconds or every 100. Bybit's order book topic pushes the top level every 10 milliseconds, fifty levels every 20, two hundred every 100 and a thousand every 200. Kraken's book channel sends an update per event with no fixed interval, and notes that one message may carry several changes to the same price level. Coinbase's level2 channel is a snapshot followed by incremental updates, again without a published cadence. Converted to a sixty-hertz frame, the fastest documented tier alone delivers more than one message per frame.

Documented order-book update cadence per venue, converted to updates per 16.7 millisecond frame Horizontal bars, updates per frame: Bybit top level at 10 milliseconds, 1.67; Bybit fifty levels at 20 milliseconds, 0.83; Binance at 100 milliseconds, 0.17; Bybit two hundred levels at 100 milliseconds, 0.17; Binance at 1,000 milliseconds, 0.02. Kraken and Coinbase are per event and shown as unbounded. A marker at 1.0 is the line above which coalescing is required. One venue can already exceed a frame Order-book updates per 60 Hz frame, from each venue's documented push interval Bybit, top level, 10 ms 1.67 Bybit, 50 levels, 20 ms 0.83 Binance, 100 ms 0.17 Bybit, 200 levels, 100 ms 0.17 Binance, 1,000 ms 0.02 Kraken, Coinbase per event, no ceiling; bursts follow the market 1.0: coalesce above this over one per frame under one per frame
Source: the WebSocket documentation of Binance, Bybit, Kraken and Coinbase, September 2026; per-frame values are the author's conversion at 16.7 ms per frame.

Two things in the chart matter more than the individual bars. The first is that the venues with no published cadence, the per-event feeds, are the ones that burst hardest, because a burst of orders is a burst of messages, and the burst arrives when the market is moving. The second is that the numbers add. Five venues at their mid tiers, each under one update per frame on its own, sum to more than one per frame together, and a terminal that renders each message has no way to know that the sum has crossed the line until the tab is already behind.

The frame quotient

The number to watch is the frame quotient: updates received per animation frame, per venue and in total. Below one, rendering per message is harmless, because there is at most one render per frame anyway. Above one, the page is doing more layout work than it can finish before the next frame is due, the work queues, and the queue grows until the tab stutters or the browser throttles the timer. The quotient is measured, not assumed: count messages arriving between two animation frame callbacks and divide.

The rule that follows has three parts. When the quotient is under one you may render per message. When it is over one you must coalesce: fold every update that arrived since the last frame into the book's state and paint once per frame. And when the age of the oldest unprocessed message exceeds two frames, stop draining deltas and ask the venue for a fresh snapshot, because you are now spending frames catching up on history the screen will never show, and a snapshot replaces the whole backlog in one message.

From WebSocket to screen with a per-frame coalescer and the snapshot escape hatch A flow: WebSocket messages arrive into a queue; a coalescer runs once per animation frame, applies every queued delta to the book state and paints once; a monitor checks the queue's age and, if the oldest message is older than two frames, discards the queue and requests a snapshot from the venue, which resets the state in one message. Fold, then paint once WebSocket deltas arrive Queue with arrival times Coalescer, per frame apply all, paint once Screen one paint a frame oldest message older than two frames? drop the queue, request a snapshot, resume
Illustrative: the pipeline as built for a terminal subscribed to several venues; the two-frame threshold is the design's choice.

Why a snapshot beats draining

The instinct when behind is to work harder: process the queue faster, skip the paint, catch up. It is the wrong instinct because the deltas describe a path the screen will never display. If two hundred updates arrived during a stall, applying all two hundred produces the same final book as one snapshot taken now, at two hundred times the cost, and every one of those frames is a frame during which new deltas keep arriving. The snapshot flip is what makes the backlog bounded: whatever happened during the stall, recovery is one message and one paint.

Queue age over time during a burst, draining deltas versus flipping to a snapshot Two lines over time. In delta mode, the age of the oldest unprocessed message rises through the burst and keeps rising after it, because the page cannot catch up. In snapshot mode, the age rises to the two-frame threshold, then drops to zero when the snapshot arrives, and stays low. A horizontal marker shows the two-frame threshold. The backlog you drain versus the backlog you replace Age of the oldest queued message in frames during a burst drain deltas snapshot flip 12 8 4 0 two-frame threshold burst begins Draining never catches up; the flip resets the age to zero in one message.
Illustrative: the shape of the two strategies during a burst; the curves are drawn to show the behaviour, not measured.

Measuring the quotient

The measurement is two counters. The message handler increments a per-venue count on every arrival and stamps the arrival time; the animation frame callback reads the counts, divides by the frames elapsed, resets, and keeps a short rolling window so that a one-frame spike does not flip the mode. The quotient during a quiet market is a fraction; during an announcement it can be several, from a single per-event venue, for a few seconds at a time. The design has to be right for those seconds, because the quiet minutes were never the problem.

Two edge cases shape the implementation more than the steady state. The first is the background tab. Browsers stop or heavily throttle animation frames when a tab is hidden, so a terminal left in a background tab for ten minutes accumulates ten minutes of deltas that no frame will ever drain; on the visibility change back to the foreground the correct move is to discard every queue and request snapshots from every venue, which is the same escape hatch triggered by a different cause. The second is the engine behind the browser. A Node.js process that fans the venues in can coalesce on a fixed tick before forwarding, which caps the browser's quotient at the tick rate at the cost of adding up to one tick of latency to what the screen shows. For a terminal, a tick of twenty or thirty milliseconds is invisible to a human and halves the worst-case work in the tab; for anything that acts on the book automatically, the tick is latency and belongs elsewhere.

What the terminal does with this

In the terminal I built, which streams books from five exchanges over WebSockets with a Node.js engine behind the browser, the frame quotient is a number on the diagnostics panel, per venue and in total, and it is the number I watch when something feels slow. The coalescer is small: each incoming delta is appended to a per-venue queue with its arrival time; a single requestAnimationFrame loop drains every queue into the book state and triggers one render; and a check at the top of the loop compares the oldest arrival time to the current frame and flips that venue into snapshot mode if the gap is over two frames. The venues that publish snapshots on request make the flip a single subscribe message; for the ones that do not, the flip is a resubscribe, which costs a round trip and is still cheaper than the backlog.

Two details took longer than the design. The first is that the book state and the rendered state have to be separate: the coalescer updates a data structure, and the render reads it, so that folding two hundred deltas costs two hundred map updates and one layout rather than two hundred layouts. The second is that the diagnostics themselves must not be rendered per message, which is the same bug in a smaller room; the quotient is computed per frame and displayed per second.

The general lesson is older than trading terminals. Any client that consumes a feed faster than it can draw has to decide whether its unit of work is the message or the frame, and the message is the wrong answer whenever the quotient can exceed one. The frame quotient gives that decision a number, the coalescer gives it a mechanism, and the snapshot flip gives it a way back when the mechanism falls behind.

WebSocketsRealtimeFrontend Performance
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS