The first thing a quantised model forgets
Quantised models recover over 99 percent of accuracy on average, and the average hides where the loss lands. A canary set from the classes that fail first should choose the format.
The case for shipping a 4-bit model is an average, and the average is excellent. Red Hat's evaluation team ran over half a million evaluations on quantised Llama 3.1 models and found that every scheme at every size recovered over 99 percent of the full-precision score on the OpenLLM v1 benchmarks, close to 99 percent on v2, and 98.9 percent on code at 4-bit. Those numbers are why teams quantise without looking further, and they are true. What they do not say is where the missing one percent lives, and the answer from the studies that looked is that it is not spread thinly across everything. It is concentrated in a few task classes, and in those classes the loss is ten times the average.
The average and the tail
The same evaluation series publishes per-model numbers, and the 8B model at 4-bit weights is the interesting one. Its model card reports 98.9 percent recovery on OpenLLM v1, 96.1 percent on v2, and 93.0 percent on Arena-Hard, the open-ended, judged benchmark closest to real conversational use; the 8-bit version of the same model sits at or above 100 percent on all three. The ordering is the point: the more the task looks like a hard, open-ended conversation, the more the 4-bit model gives up, and the smaller the model, the more it gives up.
Two other studies name the classes. Marchisio and colleagues asked how quantisation affects multilingual models and found that multi-step mathematical reasoning degrades fastest: at 4-bit with group-wise quantisation, a 35B model dropped 13.1 percent on average across languages on the multilingual maths benchmark and 17.3 percent in Chinese, while non-Latin-script languages were hurt worst in general. Their most sobering number is the gap between what automatic metrics see and what people see: a 1.7 percent average drop in Japanese on automatic tasks corresponded to a 16.0 percent drop reported by human evaluators on realistic prompts. And an April 2026 paper separates two failure modes: signal degradation, where the computation is intact but precision erodes cumulatively, and computation collapse, where components in the early layers stop working and the signal is destroyed, which is the cliff that appears below 4-bit and which no post-hoc repair fixes.
Why this fooled nobody at the helpdesk
At CRIS I tuned quantisation against latency targets for a helpdesk model, and the averages there were honest, for a reason worth stating: a helpdesk corpus in English, with short factual answers drawn from retrieved documents, is the easy case. It has no multi-step arithmetic, no non-Latin scripts, no long open-ended reasoning, and no rare identifiers that the tokeniser splits into unfamiliar pieces. The task classes that degrade first were simply absent, so the average was also the tail. That is not a reason to trust the average elsewhere. It is a description of the one kind of workload where trusting it happens to be safe, and most workloads are not that kind.
The canary set
The practice that follows is to stop evaluating the average and start evaluating the tail, deliberately, with a small fixed set of prompts drawn only from the classes that fail first. I call it the canary set: thirty to fifty prompts, chosen for fragility rather than representativeness. Multi-step arithmetic with the working shown. Long-context lookup where the answer is a specific detail far from the question. Rare identifiers, part numbers, hashes, unusual names, that must be reproduced exactly. And, for any product with users outside English, prompts in the non-Latin scripts those users write in, judged by someone who reads them, because the multilingual study's central finding is that automatic metrics miss most of that damage.
The set is run at every quantisation level under consideration, against the full-precision model as the reference, and the selection rule is a tolerance: pick the smallest format whose canary score stays within a set margin of full precision. The margin is a product decision, a few percent for a general assistant, near zero for anything that does arithmetic on the user's behalf. Because the canaries are the fragile classes, they move long before the benchmark average does, and the rule catches the format that has quietly started to fail on the users who would notice.
Building the set
The set is easy to build badly and only slightly harder to build well. The prompts should come from real traffic where there is any, sorted into the fragile classes by hand, and from written examples where there is not, and they should exclude anything the full-precision model already gets wrong, because a canary that fails at every precision measures nothing. Each prompt gets a scoring rule appropriate to its class: exact match for arithmetic results and identifiers, where a single wrong digit is a failure; a rubric applied by a person or a strong judge model for open-ended answers, with the judge's own agreement with humans checked once on a sample. Decoding is deterministic, at zero temperature, so that a change in the score is a change in the model rather than in the dice.
Two disciplines keep the set useful over time. It stays fixed across releases, so that scores are comparable and a regression is a regression; new prompts go into a candidate pool that is promoted only at a planned revision. And it stays small, thirty to fifty prompts, because the point is that it runs in minutes on every candidate build, including the ones that nobody expected to change the model, such as a runtime upgrade that quietly changed the quantisation kernels.
When a canary trips, the response is graduated. The first step is a less aggressive format for the same weights, 8-bit instead of 4, or a 4-bit scheme with smaller groups or activation-aware calibration, and the canaries are re-run. The second, informed by the two-failure-modes finding, is to leave the early layers, where computation collapse begins, at higher precision while quantising the rest, which several toolchains now support per layer. The third is to accept the larger model and find the memory elsewhere, which is where the last figure comes in.
What the canaries do not decide
The canary set answers one question, which format is safe, and it deliberately does not answer the other, whether quantisation is worth it at all. That depends on where the memory goes, and the third figure is the reminder that for an 8B model the answer changes with context length. At short contexts the weights dominate and 4-bit halves the footprint again over 8-bit. At 128k tokens the key-value cache at 16-bit is as large as the full-precision weights, and shaving the weights to 4-bit saves a quarter of the total while spending accuracy in the fragile classes. There, the better trade is often 8-bit weights and a quantised cache, which is a different decision with its own canaries.
The rule, then, has two halves. Decide whether to quantise from the memory budget at the context length you actually serve. Decide how far to quantise from the canary set, not from the benchmark average, because the average is the last number to move and the canaries are the first. The first thing a quantised model forgets is the thing your evaluation was not looking at, and the canary set is how you make sure something is.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS