The user reads at five tokens a second
For a streamed answer, decode speed only needs to beat the reader. Adults read about 238 words a minute, so speed above a floor is invisible; spend the capacity elsewhere.
The number that gets optimised in a serving stack is tokens per second, and for a streamed chat answer it is the wrong number to optimise past a point that is easy to compute. The person on the other end reads at a fixed speed. Once the model produces text faster than that, every further token per second is invisible to them, and the capacity spent producing it could have been spent on the thing they do notice, which is how long they waited before the first word appeared. When I was tuning batching for the helpdesk assistant at CRIS against a latency target, this was the reframing that made the target make sense.
How fast people read
Brysbaert's meta-analysis of reading rate, covering 190 studies and 18,573 participants, puts the average silent reading rate for adults reading English non-fiction at 238 words per minute, fiction at 260, and reading aloud at 183. Non-fiction is the right figure for a helpdesk answer or a technical explanation. At roughly 1.3 tokens per English word for a typical subword tokeniser, 238 words a minute is about four words a second, which is a little over five tokens a second.
A single stream from a modern serving stack decodes far faster than five tokens a second. That is the observation, and it is usually stated as a success. The consequence that gets missed is that the surplus is being spent on nothing. A reader who receives forty tokens a second is not reading eight times faster; they are watching text pile up below the line they are on, and the pile does not shorten the time they will spend reading it.
Two clocks
The latency a user feels in a streamed answer has two parts, and only one of them is decode speed. The first clock runs from the moment they press enter to the moment the first token appears, and it is made of queueing, the time the request waits for a slot, and prefill, the time the model spends reading the prompt. The second clock runs from the first token to the last, and it is decode: the per-token generation rate, times the length of the answer.
The reader feels the first clock entirely. Nothing is on the screen, and every hundred milliseconds is noticed. The reader feels the second clock only when it runs slower than they read, because then the text stalls under their eyes. Above reading speed, the second clock is invisible: the answer finishes arriving before they finish reading it, and it makes no difference to them whether it finished arriving one second in or five.
The reading-speed floor
So the rule is this. Set a floor on per-stream decode rate at about twice the audience's reading speed, which for English non-fiction is roughly ten tokens a second. The factor of two covers variation between readers, the fact that people skim the start of an answer faster than they read the middle, and the way tokens arrive in bursts rather than evenly. Above the floor, per-stream speed is not a metric to improve. Below it, the reader is waiting on the model and it is the metric.
Then spend everything above the floor on the first clock. In practice that means larger batches. Continuous batching stacks more concurrent requests onto the same accelerator, which raises aggregate throughput and lowers per-stream decode rate at the same time; the vLLM paper reports throughput improvements of two to four times over earlier systems at the same latency, from managing attention memory well enough to keep batches large. The reading-speed floor says how large: push the batch until the per-stream rate reaches the floor, and no further. The capacity that frees up serves more users, which shortens the queue, which shortens the first clock for everyone.
Measuring the floor for your own audience
The floor is not a universal constant, and two of its inputs are worth measuring rather than copying. The first is the tokeniser ratio. English runs at roughly 1.3 tokens per word on common subword vocabularies, but other scripts run higher, sometimes much higher, because their characters are split into more pieces. An assistant answering in Hindi or Tamil emits more tokens per word the reader reads, so the floor in tokens per second rises even though the reader's speed in words has not changed. Measure tokens per word on your own answers, per language, and set the floor per language.
The second is the reader. Brysbaert's figure is an average across adults reading prose; a support engineer scanning a known procedure reads faster, and a person reading in their second language reads slower. Where the product can afford it, the honest measurement is direct: how long the reader keeps the answer open before acting, compared with the time the stream took to finish. If the stream finished long before they acted, the floor is comfortably met and the decode budget can be cut further.
What the floor does to the operations dashboard is the practical payoff. Aggregate tokens per second stops being the headline. In its place go two service objectives: a per-stream decode rate that stays above the floor at the 95th percentile, and a time to first token that is as low as the freed capacity can make it. The first is a constraint. The second is the thing to improve.
Where the floor does not apply
The rule is about people reading a stream, and it fails wherever that description does not hold. A model whose output feeds another program, a parser, a tool call, a second model, has no reader, and there the per-token rate is the whole latency and should be as high as the hardware allows. A model producing code that the user will scroll through rather than read in order is somewhere in between. And a user who has stopped reading and is waiting for the end, because the answer is a long list they only want the bottom of, has become a machine consumer for the duration.
It also fails for very short answers, where the first clock dominates whatever the decode rate, and for the pathological case where prefill is so long that the first token arrives after the reader has given up. Neither of those is fixed by decode speed either. They are fixed by shorter prompts and shorter queues, which is where the freed capacity goes.
What it changed
For a helpdesk answer read by a person on a screen, the floor applied cleanly, and the number that mattered was the first clock. The work that improved the experience was not faster decode. It was keeping the retrieved passages short so prefill was short, and keeping the batch large so the queue was short, with the per-stream rate held at the floor rather than maximised. Five tokens a second is the reader's speed. The model only has to beat it.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS