18 December 2025 · 5 min read

A LoRA adapter is a poor place for facts

LoRA updates are low-rank and learn less new knowledge than full fine-tuning. Adapters hold format well and facts badly, and facts change faster than anyone retrains.

At CRIS I did both halves of the usual argument on one system: I fine-tuned open models with parameter-efficient adapters for a helpdesk assistant, and I built the retrieval pipeline that fed the same assistant its documents. The two halves are usually presented as competitors, fine-tuning against retrieval, and the decision between them is usually made by taste. Doing both on one problem made the division of labour obvious, and the division turned out to depend on how fast each kind of knowledge decays, measured against how often anyone will actually retrain, rather than on which technique is better.

What a low-rank update can hold

A LoRA adapter does not change a model's weights. It adds, to selected weight matrices, a correction that is the product of two thin matrices, so that the correction has a rank no higher than their shared inner dimension, r. The number of trainable parameters is therefore small and easy to compute from the architecture. For an eight-billion-parameter model of the Llama 3 family, whose published architecture has 32 layers, a hidden size of 4,096, grouped key and value projections of 1,024, and a feed-forward width of 14,336, adapting only the query and value projections adds about 426,000 parameters per unit of rank; adapting every linear layer in each block adds about 2.6 million per unit of rank.

Trainable parameters in a LoRA adapter against rank, for an 8B Llama-family model Two rising lines on a logarithmic axis. Adapting query and value projections only: 1.7 million parameters at rank 4, 3.4 million at 8, 6.8 million at 16, 13.6 million at 32, 27 million at 64. Adapting all linear layers: 10.5 million at rank 4, 21 million at 8, 42 million at 16, 84 million at 32, 168 million at 64. Even the largest is about two percent of the eight billion base parameters. Small by construction Trainable parameters, derived from the 8B architecture; log scale all linear layers q and v only 1B 100M 10M 1M 4 8 16 32 64 LoRA rank, r 168M, about 2% of the base 27M
Source: derived by the author from the layer dimensions in the Llama 3 paper; parameters per unit of rank are the sums of input and output dimensions of the adapted matrices, times 32 layers.

Two percent of the parameters, at a rank most people never use, and the correction is constrained to a low-dimensional subspace within that. That is a lot of capacity for shaping how a model responds. It is not a lot of capacity for storing what the model knows, and the empirical work agrees. Biderman and colleagues' LoRA Learns Less and Forgets Less compared adapters with full fine-tuning across code and mathematics, and found that LoRA learns substantially less of the new domain than full fine-tuning does, while also forgetting less of what the base model could do. The trade is the mechanism: a low-rank update cannot rewrite the model's knowledge, so it preserves the old and absorbs less of the new.

What the paper measured

The result deserves a closer look, because it is the empirical anchor for everything that follows. The study compared LoRA against full fine-tuning on two target domains, programming and mathematics, in two regimes each: continued pretraining on large unlabelled corpora and instruction fine-tuning on smaller labelled sets. Learning was measured on the target domain's benchmarks, and forgetting was measured on general tasks the base model already did well, so that each run produced a point on a learning-versus-forgetting plane. Across the regimes, LoRA sat consistently below full fine-tuning on the target domain and above it on the retained general ability. The gap on learning was largest in continued pretraining on code, which is the regime closest to "teach the model a body of facts", and smallest in instruction fine-tuning on maths, which is closest to "teach the model a way of answering".

The paper also looked at why, and the explanation matches the arithmetic above. When the authors examined the weight changes that full fine-tuning produced, they found perturbations whose rank was ten to a hundred times higher than the ranks LoRA is normally run at. Full fine-tuning is not using a low-rank update that LoRA could match with a bigger r; it is using a genuinely high-rank one, and a low-rank adapter is a projection of it. LoRA, in return, acts as a stronger regulariser than the usual alternatives, keeping the model's outputs closer to the base and its generations more diverse. That is the behaviour-preserving property that makes an adapter safe to ship for the things it is good at.

Behaviour versus facts

That result reads as a limitation, and for one class of knowledge it is a feature. The things a helpdesk assistant most needs to learn from fine-tuning are behavioural: answer in this format, refuse these categories, cite in this style, keep this tone, stop after the answer. Those are patterns, not facts, and patterns live comfortably in a low-rank correction to how attention and projections behave. An adapter trained on a few thousand well-formed examples teaches them reliably, and forgetting less of the base model is exactly what you want while it does.

The things the assistant most needs to know are facts: which form to file, which screen to open, what the policy says this quarter, which phone number moved. Facts are the wrong shape for a low-rank update, they are the category the paper shows adapters learning least of, and they have a second problem that is worse than capacity. They change.

The half-life split

So the rule I use is a question about decay rather than about technique. For each class of knowledge the system needs, estimate how long until half the answers in that class have changed. Then compare that half-life with the retraining cadence, how often anyone will actually build, evaluate and ship a new adapter, which for most teams is quarterly at best and, honestly, less often than that.

Where each class of knowledge belongs, by decay and kind A two by two grid. Horizontal axis: kind of knowledge, behavioural on the left and factual on the right. Vertical axis: half-life, shorter than the retraining cadence at the bottom, longer at the top. Bottom right: fast-changing facts, retrieval. Top right: stable facts, the one contested cell, where rank decides. Top left: stable behaviour, the adapter. Bottom left: fast-changing behaviour, prompts and retrieved instructions. The adapter format, tone, refusals, style Contested: rank decides stable facts, domain vocabulary Prompt and instructions behaviour that changes with policy Retrieval procedures, numbers, this quarter's rules Kind of knowledge: behavioural to factual Half-life: shorter than retraining to longer
Illustrative: the split as a grid; the dot is where most of a helpdesk's knowledge sits.

Anything with a half-life shorter than the retraining cadence goes to retrieval, whatever its kind, because a fact baked into an adapter is wrong from the day it changes until the day someone retrains, and that gap is the half-life's whole length. Anything with a long half-life and a behavioural kind goes to the adapter, because that is what a low-rank update is good at and the knowledge will not rot before the next release. The one contested cell is long-lived factual knowledge, domain vocabulary, the stable structure of a product, the things that were true five years ago and will be true in five more, and there the adapter and retrieval genuinely compete. Rank decides it: a higher rank absorbs more, at the cost of forgetting more, and the paper's curves are the guide to where that trade sits for a given domain.

Estimating the half-life

The half-life is easier to estimate than it sounds, because the organisation already knows it. Ask how often each kind of document is revised.

Pricing changes monthly. Procedures change with each release. Policy changes when the regulator does. Product names change at the whim of marketing. The tone of a good answer changes never. Put those on one axis and the retraining cadence on the same axis, and the classes sort themselves.

Knowledge classes by half-life against a quarterly retraining cadence Horizontal bars on a logarithmic scale of weeks: pricing about 4 weeks, procedures about 12, product names about 26, policy about 52, tone and format effectively unbounded. A vertical marker at 13 weeks is the retraining cadence. Everything left of it belongs in retrieval. Which classes decay before the next retrain Half-life in weeks, log scale; example values for a helpdesk, stated as assumptions Pricing 4 Procedures 12 Product names 26 Policy 52 Tone and format effectively never retraining cadence: 13 weeks retrieval contested adapter
Illustrative: half-lives are example values for a helpdesk domain, chosen to show the sort against a quarterly cadence; a real estimate comes from the organisation's own revision history.

What this meant in practice

The split settled a question that had been consuming effort in the wrong place. The adapter's training set stopped being a dump of the manuals and became a few thousand examples of good answers: the right format, the right refusals, the right citation style, the right way to say that a question needs a human. The manuals went into the retrieval index, where a revised procedure replaced the old one the same afternoon, with no training run. And the evaluation split the same way: the adapter was judged on whether answers had the right shape, and retrieval was judged on whether the right passage was found, which are different failures with different fixes.

The rule generalises past helpdesks. Any time the question is "fine-tune or retrieve", replace it with two questions: what kind of knowledge is this, and how long until half of it is wrong. The adapter is where you put the things that will still be true at the next release. Everything else is a document, and documents are for retrieving.

Fine-TuningLoRARAG
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS