Four bits are enough, three are not
Quantisation error is a group's range divided by its levels, and a few outlier weights set the range for everyone. Sixteen levels leave the bulk enough resolution; eight do not.
Everyone who serves language models learns the same rule of thumb: 8-bit weights are free, 4-bit weights are nearly free, and below 4 bits the model gets noticeably worse, fast. The rule is right, and the explanations offered for it are usually hand-waving about information content. The actual reason is arithmetic that fits on a napkin.
Weight quantisation is rounding. The rounding error is the range of each group of weights divided by the number of levels available, and the range is set by the largest weight in the group, which is an outlier that almost nothing else in the group resembles. Sixteen levels leave enough of the range for the bulk of the weights after the outlier has taken its share. Eight do not. Everything published below 4 bits is, at bottom, a way of stopping the outlier from taxing everyone else.
The cliff, measured
The cliff is visible in any careful evaluation, and the cleanest public one I know is the table that accompanied the introduction of the k-quant formats in llama.cpp, which reports perplexity for the 7-billion-parameter LLaMA at each format alongside the file size. Against a 16-bit baseline of 5.9066, the 4-bit medium format lands at 5.9601, under one percent worse, for a file 3.4 times smaller. The 3-bit medium format is 6.1503, four percent worse, and the 3-bit small format is 6.4571, over nine percent worse. The 2-bit format is 6.7764, fifteen percent worse, and its file is barely smaller than the 3-bit small one. The curve is flat from 16 bits down to about 4.5 and then bends sharply.
What a weight matrix looks like
To see why the cliff is there, look at the weights themselves. I took a public checkpoint small enough to download over lunch, the 135-million-parameter SmolLM2, and pulled one matrix from the middle of the network, the down-projection of a feed-forward block, 576 by 1,536, about 885,000 weights. Measured in units of the matrix's own standard deviation, 99.45 percent of the weights lie within three sigma and 99.94 percent within four. Sixty-two weights, seven in a hundred thousand, lie beyond six sigma. The largest is at 18.8 sigma. The distribution is a near-Gaussian bulk with a tail that is thin, long, and decisive.
The range tax
Now quantise it. A uniform quantiser with b bits has 2 to the b levels spread evenly across a range, and the simplest range is symmetric around zero out to the largest absolute weight in the group. The step between levels is twice the largest weight divided by the number of levels minus one, and the rounding error is, on average, a quarter of a step. Everything about the cliff follows from what "the group" is.
If the group is the whole matrix, the largest weight is 18.8 sigma, and the 4-bit step is 2.5 sigma. Almost every weight in the matrix rounds to one of three levels, and the model is destroyed. Eighty percent of the quantiser's range lies beyond the 99.9th percentile of the weights, spent on values that fewer than one weight in a thousand ever takes. That share is what I call the range tax: the fraction of the range, and so of the levels, that the outliers take from everyone else.
The fix everyone uses is to shrink the group. With groups of 128 consecutive weights, the median group's largest value is 2.93 sigma, the 4-bit step is 0.39 sigma, and ten of the sixteen levels fall inside plus or minus two sigma where the bulk lives. The mean rounding error, measured, is 0.10 sigma. Only 5.5 percent of groups contain a weight beyond four sigma, so the tax is paid in a few groups rather than everywhere, and it averages about one percent of the range.
Why the cliff is between four and three
Hold the group at 128 and drop to 3 bits. The same median range of 2.93 sigma is now split into eight levels, the step is 0.84 sigma, and only four levels fall inside plus or minus two sigma. The measured rounding error doubles, to 0.21 sigma. At 2 bits the step is 1.95 sigma, two levels serve the bulk, and the error is 0.51 sigma, half the width of the distribution. The perplexity table above is that arithmetic seen from the outside: the one-percent penalty at 4 bits, the four-to-nine-percent penalty at 3 bits, the fifteen at 2.
The reason the threshold is where it is comes from the shape of the bulk. A near-Gaussian distribution needs a step of a few tenths of a sigma before rounding noise is small relative to the signal in each weight, and the smallest range a group of a hundred-odd samples can have is about three sigma either side, because the largest of a hundred draws from a Gaussian sits near there even with no outliers at all. Six sigma of range at a step of 0.4 needs about sixteen levels. That is four bits. Eight levels over the same range is a step of 0.84 sigma, and the noise is no longer small. The cliff is not a property of any model; it is a property of a Gaussian bulk with a hundred-sample range, and every model whose weights look like that will have it in the same place.
Every sub-4-bit method is an outlier method
Seen this way, the research below 4 bits is a set of ways to lower the range tax, and they are three, not thirty. The first is group size: smaller groups mean each outlier taxes fewer neighbours, and in the matrix above, going from groups of 1,024 to 128 to 32 takes the share of groups with a weight past four sigma from 28 percent to 5.5 to 1.7, and the average tax from about six percent of the range to one to under half a percent, at the cost of storing a scale per group. The second is rotation: multiply the weights by an orthogonal matrix before quantising, as QuaRot and its successors do, so that a single large entry is spread across many coordinates and the largest value in every group shrinks toward the Gaussian expectation, which lowers the range without changing what the layer computes. The third is mixed precision: find the columns that carry the outliers, which activation-aware methods identify from the activations they multiply, and keep those at higher precision while the rest go to 4 bits or below.
A fourth family, the non-uniform and importance-weighted codebooks that the k-quants themselves belong to, places the levels where the weights actually are rather than evenly across the range, which is a way of refusing to spend levels on the empty part of it; it lowers the tax without shrinking the range, and it composes with the other three.
Each of the four attacks the same quantity. None of them adds resolution to the bulk by fiat; they stop the bulk from paying for the tail. That is why the honest description of a 3-bit or 2-bit result is never "the model tolerates fewer bits" and always "the method found a way to make the range smaller", and it is why, when I choose between 8-bit and 4-bit for a served model, the question I actually ask is how the quantiser handles outliers, not how many bits it claims.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS