15 February 2025 · 5 min read

Four bits are enough, three are not

Quantisation error is a group's range divided by its levels, and a few outlier weights set the range for everyone. Sixteen levels leave the bulk enough resolution; eight do not.

Everyone who serves language models learns the same rule of thumb: 8-bit weights are free, 4-bit weights are nearly free, and below 4 bits the model gets noticeably worse, fast. The rule is right, and the explanations offered for it are usually hand-waving about information content. The actual reason is arithmetic that fits on a napkin.

Weight quantisation is rounding. The rounding error is the range of each group of weights divided by the number of levels available, and the range is set by the largest weight in the group, which is an outlier that almost nothing else in the group resembles. Sixteen levels leave enough of the range for the bulk of the weights after the outlier has taken its share. Eight do not. Everything published below 4 bits is, at bottom, a way of stopping the outlier from taxing everyone else.

The cliff, measured

The cliff is visible in any careful evaluation, and the cleanest public one I know is the table that accompanied the introduction of the k-quant formats in llama.cpp, which reports perplexity for the 7-billion-parameter LLaMA at each format alongside the file size. Against a 16-bit baseline of 5.9066, the 4-bit medium format lands at 5.9601, under one percent worse, for a file 3.4 times smaller. The 3-bit medium format is 6.1503, four percent worse, and the 3-bit small format is 6.4571, over nine percent worse. The 2-bit format is 6.7764, fifteen percent worse, and its file is barely smaller than the 3-bit small one. The curve is flat from 16 bits down to about 4.5 and then bends sharply.

Perplexity against model file size for the 7B LLaMA k-quant formats A scatter with a connecting line. Points from right to left: Q6_K at 5.15 gigabytes and perplexity 5.911; Q5_K_M at 4.45 and 5.921; Q5_K_S at 4.33 and 5.942; Q4_K_M at 3.80 and 5.960; Q4_K_S at 3.56 and 6.022; Q3_K_M at 3.06 and 6.150; Q3_K_S at 2.75 and 6.457; Q2_K at 2.67 and 6.776. A dashed reference line marks the 16-bit perplexity of 5.907. The curve is flat down to about 3.8 gigabytes and rises steeply below. Flat to four bits, then a cliff Perplexity of LLaMA 7B by quantised file size in gigabytes; 16-bit reference 5.907 6.8 6.6 6.4 6.2 6.0 5.8 3 GB 4 GB 5 GB 16-bit: 5.907 Q2_K, 6.776 Q3_K_S, 6.457 Q3_K_M, 6.150 Q4_K_S Q4_K_M, 5.960 Q5_K Q6_K within one percent of 16-bit the cliff: 4 to 15 percent worse
Source: the perplexity and file-size table in the k-quants pull request to llama.cpp, June 2023, for the 7B model; positions are drawn from the table's values.

What a weight matrix looks like

To see why the cliff is there, look at the weights themselves. I took a public checkpoint small enough to download over lunch, the 135-million-parameter SmolLM2, and pulled one matrix from the middle of the network, the down-projection of a feed-forward block, 576 by 1,536, about 885,000 weights. Measured in units of the matrix's own standard deviation, 99.45 percent of the weights lie within three sigma and 99.94 percent within four. Sixty-two weights, seven in a hundred thousand, lie beyond six sigma. The largest is at 18.8 sigma. The distribution is a near-Gaussian bulk with a tail that is thin, long, and decisive.

Histogram of one weight matrix in units of its standard deviation, with the 4-bit levels for a typical group drawn beneath Twenty bars on a logarithmic count axis, one per sigma from minus ten to plus ten. The central two bins hold about 311,000 weights each; the bins at plus and minus three to four hold about 2,100; at five to six, a few dozen; beyond eight, a handful, with the largest weight at 18.8 sigma off the chart. Below the axis, sixteen gold ticks mark the levels a 4-bit quantiser has for a typical group of 128 weights whose largest value is 2.93 sigma; the ticks are 0.39 sigma apart. A near-Gaussian bulk and a thin, long tail Weights per one-sigma bin, log scale; SmolLM2-135M, layer 22 down-projection, 884,736 weights 1M 10k 100 1 -10 -5 0 5 10 4-bit levels, 0.39 sigma apart Gold bars: 99.45 percent within three sigma; the largest weight is at 18.8 sigma.
Source: computed by the author from the public SmolLM2-135M checkpoint; bars are counts per one-sigma bin, the ticks are the levels of a symmetric 4-bit quantiser for a group of 128 whose largest weight is at the median, 2.93 sigma.

The range tax

Now quantise it. A uniform quantiser with b bits has 2 to the b levels spread evenly across a range, and the simplest range is symmetric around zero out to the largest absolute weight in the group. The step between levels is twice the largest weight divided by the number of levels minus one, and the rounding error is, on average, a quarter of a step. Everything about the cliff follows from what "the group" is.

If the group is the whole matrix, the largest weight is 18.8 sigma, and the 4-bit step is 2.5 sigma. Almost every weight in the matrix rounds to one of three levels, and the model is destroyed. Eighty percent of the quantiser's range lies beyond the 99.9th percentile of the weights, spent on values that fewer than one weight in a thousand ever takes. That share is what I call the range tax: the fraction of the range, and so of the levels, that the outliers take from everyone else.

The fix everyone uses is to shrink the group. With groups of 128 consecutive weights, the median group's largest value is 2.93 sigma, the 4-bit step is 0.39 sigma, and ten of the sixteen levels fall inside plus or minus two sigma where the bulk lives. The mean rounding error, measured, is 0.10 sigma. Only 5.5 percent of groups contain a weight beyond four sigma, so the tax is paid in a few groups rather than everywhere, and it averages about one percent of the range.

Per-group scaling: the range, the step, and one outlier stretching the step for the whole group Two number lines. Top, a group with no outlier: the weights cluster within about three sigma, the sixteen levels span that cluster, and the step is small. Bottom, a group with one weight at eight sigma: the same sixteen levels now span eight sigma, the step is nearly three times larger, and most of the levels lie in empty space between the cluster and the outlier. One weight sets the step for 127 others GROUP WITHOUT AN OUTLIER: range 3 sigma, step 0.4 sigma 128 weights, within 3 sigma 16 levels in the cluster GROUP WITH ONE OUTLIER AT 8 SIGMA: range 8 sigma, step 1.07 sigma the same 128 weights one weight at 8 sigma Same 16 levels, stretched over the outlier: six inside the cluster, ten in empty space. the range tax is the share of levels in the empty space, here over sixty percent
Illustrative: two groups drawn to scale on a sigma axis; the step sizes are computed for a symmetric 16-level quantiser over each range.

Why the cliff is between four and three

Hold the group at 128 and drop to 3 bits. The same median range of 2.93 sigma is now split into eight levels, the step is 0.84 sigma, and only four levels fall inside plus or minus two sigma. The measured rounding error doubles, to 0.21 sigma. At 2 bits the step is 1.95 sigma, two levels serve the bulk, and the error is 0.51 sigma, half the width of the distribution. The perplexity table above is that arithmetic seen from the outside: the one-percent penalty at 4 bits, the four-to-nine-percent penalty at 3 bits, the fifteen at 2.

The reason the threshold is where it is comes from the shape of the bulk. A near-Gaussian distribution needs a step of a few tenths of a sigma before rounding noise is small relative to the signal in each weight, and the smallest range a group of a hundred-odd samples can have is about three sigma either side, because the largest of a hundred draws from a Gaussian sits near there even with no outliers at all. Six sigma of range at a step of 0.4 needs about sixteen levels. That is four bits. Eight levels over the same range is a step of 0.84 sigma, and the noise is no longer small. The cliff is not a property of any model; it is a property of a Gaussian bulk with a hundred-sample range, and every model whose weights look like that will have it in the same place.

Every sub-4-bit method is an outlier method

Seen this way, the research below 4 bits is a set of ways to lower the range tax, and they are three, not thirty. The first is group size: smaller groups mean each outlier taxes fewer neighbours, and in the matrix above, going from groups of 1,024 to 128 to 32 takes the share of groups with a weight past four sigma from 28 percent to 5.5 to 1.7, and the average tax from about six percent of the range to one to under half a percent, at the cost of storing a scale per group. The second is rotation: multiply the weights by an orthogonal matrix before quantising, as QuaRot and its successors do, so that a single large entry is spread across many coordinates and the largest value in every group shrinks toward the Gaussian expectation, which lowers the range without changing what the layer computes. The third is mixed precision: find the columns that carry the outliers, which activation-aware methods identify from the activations they multiply, and keep those at higher precision while the rest go to 4 bits or below.

A fourth family, the non-uniform and importance-weighted codebooks that the k-quants themselves belong to, places the levels where the weights actually are rather than evenly across the range, which is a way of refusing to spend levels on the empty part of it; it lowers the tax without shrinking the range, and it composes with the other three.

Each of the four attacks the same quantity. None of them adds resolution to the bulk by fiat; they stop the bulk from paying for the tail. That is why the honest description of a 3-bit or 2-bit result is never "the model tolerates fewer bits" and always "the method found a way to make the range smaller", and it is why, when I choose between 8-bit and 4-bit for a served model, the question I actually ask is how the quantiser handles outliers, not how many bits it claims.

QuantisationLinear AlgebraInference
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS