13 January 2026 · 5 min read

The user reads at five tokens a second

For a streamed answer, decode speed only needs to beat the reader. Adults read about 238 words a minute, so speed above a floor is invisible; spend the capacity elsewhere.

The number that gets optimised in a serving stack is tokens per second, and for a streamed chat answer it is the wrong number to optimise past a point that is easy to compute. The person on the other end reads at a fixed speed. Once the model produces text faster than that, every further token per second is invisible to them, and the capacity spent producing it could have been spent on the thing they do notice, which is how long they waited before the first word appeared. When I was tuning batching for the helpdesk assistant at CRIS against a latency target, this was the reframing that made the target make sense.

How fast people read

Brysbaert's meta-analysis of reading rate, covering 190 studies and 18,573 participants, puts the average silent reading rate for adults reading English non-fiction at 238 words per minute, fiction at 260, and reading aloud at 183. Non-fiction is the right figure for a helpdesk answer or a technical explanation. At roughly 1.3 tokens per English word for a typical subword tokeniser, 238 words a minute is about four words a second, which is a little over five tokens a second.

Human reading rates converted to tokens per second Horizontal bars: reading aloud about 4.0 tokens per second, silent non-fiction about 5.2, silent fiction about 5.6. A marker at about 10 tokens per second shows the reading-speed floor, twice the non-fiction rate, that this post proposes for a streamed answer. How fast the reader can go Tokens per second, at 1.3 tokens per word Reading aloud, 183 wpm 4.0 Silent non-fiction, 238 wpm 5.2 Silent fiction, 260 wpm 5.6 floor: 10.4, twice non-fiction A stream above the floor is never the thing the reader is waiting on.
Source: reading rates from Brysbaert (2019); the tokens-per-word factor is a stated assumption, and the floor is this post's rule.

A single stream from a modern serving stack decodes far faster than five tokens a second. That is the observation, and it is usually stated as a success. The consequence that gets missed is that the surplus is being spent on nothing. A reader who receives forty tokens a second is not reading eight times faster; they are watching text pile up below the line they are on, and the pile does not shorten the time they will spend reading it.

Two clocks

The latency a user feels in a streamed answer has two parts, and only one of them is decode speed. The first clock runs from the moment they press enter to the moment the first token appears, and it is made of queueing, the time the request waits for a slot, and prefill, the time the model spends reading the prompt. The second clock runs from the first token to the last, and it is decode: the per-token generation rate, times the length of the answer.

One streamed request as a timeline with two clocks A horizontal timeline. From the request to the first token: a queueing segment then a prefill segment, together labelled time to first token, which the reader experiences as waiting. From the first token to the last: a long decode segment, which the reader experiences as reading, as long as tokens arrive faster than they are read. What the reader is waiting for The first clock is felt; the second is hidden above the floor waiting reading queue prefill decode, one token at a time time to first token: waiting time per output token: reading, if above the floor Capacity moved from the grey segment to the brick one is capacity the reader feels.
Illustrative: the two clocks of a streamed response; segment lengths are drawn for legibility, not to scale.

The reader feels the first clock entirely. Nothing is on the screen, and every hundred milliseconds is noticed. The reader feels the second clock only when it runs slower than they read, because then the text stalls under their eyes. Above reading speed, the second clock is invisible: the answer finishes arriving before they finish reading it, and it makes no difference to them whether it finished arriving one second in or five.

The reading-speed floor

So the rule is this. Set a floor on per-stream decode rate at about twice the audience's reading speed, which for English non-fiction is roughly ten tokens a second. The factor of two covers variation between readers, the fact that people skim the start of an answer faster than they read the middle, and the way tokens arrive in bursts rather than evenly. Above the floor, per-stream speed is not a metric to improve. Below it, the reader is waiting on the model and it is the metric.

Then spend everything above the floor on the first clock. In practice that means larger batches. Continuous batching stacks more concurrent requests onto the same accelerator, which raises aggregate throughput and lowers per-stream decode rate at the same time; the vLLM paper reports throughput improvements of two to four times over earlier systems at the same latency, from managing attention memory well enough to keep batches large. The reading-speed floor says how large: push the batch until the per-stream rate reaches the floor, and no further. The capacity that frees up serves more users, which shortens the queue, which shortens the first clock for everyone.

Per-stream decode rate and aggregate throughput against batch size, a model Two curves over batch sizes from 1 to 64. Per-stream tokens per second falls from 60 at batch 1 to about 6 at batch 64. Aggregate tokens per second rises from 60 to about 380. A horizontal marker at 10 tokens per second, the reading-speed floor, crosses the per-stream curve near batch 32, which is where the model says to stop growing the batch. Push the batch to the floor, then stop Tokens per second, from a stated model; log scale per stream aggregate 1,000 100 10 1 1 2 4 8 16 32 64 Concurrent requests in the batch reading-speed floor, 10 tokens per second stop near 32
Illustrative: constructed from a model in which a step of decode takes 16 ms plus 1.2 ms per request in the batch, chosen to show the shape of the trade rather than to describe any hardware.

Measuring the floor for your own audience

The floor is not a universal constant, and two of its inputs are worth measuring rather than copying. The first is the tokeniser ratio. English runs at roughly 1.3 tokens per word on common subword vocabularies, but other scripts run higher, sometimes much higher, because their characters are split into more pieces. An assistant answering in Hindi or Tamil emits more tokens per word the reader reads, so the floor in tokens per second rises even though the reader's speed in words has not changed. Measure tokens per word on your own answers, per language, and set the floor per language.

The second is the reader. Brysbaert's figure is an average across adults reading prose; a support engineer scanning a known procedure reads faster, and a person reading in their second language reads slower. Where the product can afford it, the honest measurement is direct: how long the reader keeps the answer open before acting, compared with the time the stream took to finish. If the stream finished long before they acted, the floor is comfortably met and the decode budget can be cut further.

What the floor does to the operations dashboard is the practical payoff. Aggregate tokens per second stops being the headline. In its place go two service objectives: a per-stream decode rate that stays above the floor at the 95th percentile, and a time to first token that is as low as the freed capacity can make it. The first is a constraint. The second is the thing to improve.

Where the floor does not apply

The rule is about people reading a stream, and it fails wherever that description does not hold. A model whose output feeds another program, a parser, a tool call, a second model, has no reader, and there the per-token rate is the whole latency and should be as high as the hardware allows. A model producing code that the user will scroll through rather than read in order is somewhere in between. And a user who has stopped reading and is waiting for the end, because the answer is a long list they only want the bottom of, has become a machine consumer for the duration.

It also fails for very short answers, where the first clock dominates whatever the decode rate, and for the pathological case where prefill is so long that the first token arrives after the reader has given up. Neither of those is fixed by decode speed either. They are fixed by shorter prompts and shorter queues, which is where the freed capacity goes.

What it changed

For a helpdesk answer read by a person on a screen, the floor applied cleanly, and the number that mattered was the first clock. The work that improved the experience was not faster decode. It was keeping the retrieved passages short so prefill was short, and keeping the batch large so the queue was short, with the per-stream rate held at the floor rather than maximised. Five tokens a second is the reader's speed. The model only has to beat it.

InferenceLatencyLLM Serving
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS