24 March 2026 · 6 min read

Fine-tuning taught it to sound sure

Fine-tuning makes a model more accurate and less calibrated at once: cross-entropy keeps rewarding larger logits after the answers stop changing. Refit the temperature last.

A fine-tuned model is usually more accurate than the base model on the task it was tuned for, and it is usually more confident, and the second effect outruns the first. The reason is not mysterious. Cross-entropy loss on labels that are nearly deterministic keeps rewarding larger logits long after the argmax has stopped changing, so the weights keep growing in the direction that makes the model's answers sharper without making them righter. The result is a model that sounds sure in proportion to how long it was trained rather than how often it is correct. The fix is one scalar, fitted on held-out data after the final fine-tune, and the rule that matters is the word "after": no temperature survives a change to the weights, so it is the last step of every training run or it is not there at all.

Watching the logits grow

The mechanism is easiest to see in a model small enough to watch. I trained a logistic regression, the simplest model that has logits, on eighty examples of a two-class problem in twenty dimensions with plain gradient descent on cross-entropy, no regularisation, and recorded four things at intervals: training accuracy, accuracy on four thousand held-out examples, the mean confidence the model reported on those held-out examples, and the norm of the weight vector, which is the size of the logits.

Weight norm and mean confidence keep rising after held-out accuracy has stopped changing, in a toy model Two lines against training epochs on a log scale from 1 to 3,000. Held-out accuracy is flat at about 90 percent from epoch 10 onward. Mean reported confidence rises from 63 percent at epoch 1 to 97 percent at epoch 3,000. A third series, the weight norm, rises from 0.4 to 9.3 over the same range, shown as labelled points. The answers stop changing at epoch ten; the confidence does not Toy logistic model, 80 train examples, log-scale epochs mean confidence held-out accuracy 100% 90% 80% 70% 60% 1 10 100 1,000 63% 97% 90% weight norm: 0.4 1.5 3.5 7.1 9.3 Training accuracy hit 100 percent at epoch 10; everything after is the logits growing.
Source: computed by the author from a toy logistic regression trained with gradient descent on cross-entropy, following the calibration measures of Guo et al.; the construction is described in the text and the numbers are from one run.

By epoch ten the model fits its eighty training examples perfectly and its held-out accuracy has settled at about 90 percent, where it stays. The weight norm keeps climbing, from 1.5 at epoch ten to 9.3 at epoch three thousand, and the mean confidence climbs with it, from 83 percent to 97 percent, on a model that is right 90 percent of the time. The loss is still going down through all of that, because on training examples that are already classified correctly, a larger logit gives a smaller cross-entropy. The optimiser is doing exactly what it was asked. What it was asked has nothing to do with the held-out data.

The same thing happens in a language model fine-tuned on a small labelled set, at a scale where nobody can watch the weight norm, and the published evidence says so. The GPT-4 technical report shows the pre-trained model well calibrated on a multiple-choice benchmark and the post-trained model markedly less so, with the plain statement that post-training hurts calibration. The 2026 HypeLoRA work on low-rank adapters treats calibration as a first-class metric precisely because adapted models trade it away for task accuracy, and finds that constraining the adaptation acts as a regulariser that improves calibration at a cost in task performance. Larger logits after the argmax has settled is the mechanism in both.

One scalar, fitted last

The repair is old and almost embarrassingly small. Guo and colleagues showed in On Calibration of Modern Neural Networks that dividing the logits by a single temperature, fitted on held-out data to minimise the negative log-likelihood, restores calibration without changing a single prediction, because dividing every logit by the same positive number does not change which is largest. On the toy model the fitted temperature is 3.9; the expected calibration error on the held-out set falls from 0.072 to 0.013, the accuracy stays at 90.1 percent, and the mean confidence comes down from 97 percent to 89, which is what a model right nine times in ten should report.

Reliability diagram of the toy model before and after temperature scaling Confidence on the horizontal axis against observed accuracy on the vertical, with a diagonal for perfect calibration. Before scaling, the points sit below the diagonal: confidence 0.65 with accuracy 0.48, 0.75 with 0.54, 0.85 with 0.64, and the large top bin at confidence 0.996 with accuracy 0.93. After scaling with a fitted temperature of 3.9, the points sit on the diagonal: 0.65 with 0.67, 0.75 with 0.77, 0.86 with 0.87, 0.97 with 0.98. Before: sure and wrong. After: sure in proportion Observed against reported confidence, held-out, 0.1 bins before, ECE 0.072 after, ECE 0.013 1.0 0.875 0.75 0.625 0.5 0.5 0.75 1.0 Reported confidence perfect calibration below the line: too sure
Source: computed by the author from the toy model after 3,000 epochs, with a temperature of 3.9 fitted on a separate validation set by minimising negative log-likelihood, the procedure of Guo et al.; bins with fewer than twenty examples are omitted.

Two properties of the fit matter for the rule. The first is that it must be fitted on data the training never saw, because fitting it on the training set finds a temperature near one: the model is, after all, perfectly right on those examples, and perfectly right deserves perfect confidence. The second is that the temperature is a property of a specific set of weights. Change the weights, by another epoch, by a second fine-tune, by merging an adapter, by quantising, and the logits change scale, and the temperature that was right for the old logits is wrong for the new ones. There is no such thing as a calibrated model; there is a calibrated checkpoint.

The temperature-last rule

So the rule is about pipeline order. The last step of every training run, after the final weight update of any kind, is to refit the temperature on held-out data and to record the reliability diagram beside the accuracy. A checkpoint shipped without both has not been evaluated; it has been scored, which is half of an evaluation. And any step that touches the weights afterwards, however small, invalidates the fit and sends the checkpoint back to the calibration stage.

The training pipeline with the temperature fit as the mandatory last stage, and an invalidation arrow from any weight change A flow: base model, fine-tune, evaluate accuracy, fit temperature on held-out data and record the reliability diagram, ship. An arrow from a box labelled any weight change, including merging, quantising or further tuning, loops back to the temperature fit, marking it invalid. Calibrate last, and again after anything Base model Fine-tune Accuracyon held-out Fit temperature on held-out data record reliability Ship Any weight change merge, quantise, tune again invalidates the fit A checkpoint without a reliability diagram has been scored, not evaluated.
Illustrative: the pipeline order the rule enforces; the loop from the lower box is the part most pipelines lack.

What the rule costs, and what it does not

The objection to the rule is the held-out set: fitting the temperature needs labelled examples the training never touched, and for a fine-tune on a small labelled set those examples are precious. In practice the cost is small, because a single scalar is fitted, not a model, and a few hundred examples pin it down; the toy above used two thousand and would have been fine with a tenth of that. What the rule does cost is a place in the pipeline, a step that runs after the merge and before the export, and a record, the reliability diagram, that has to be stored beside the accuracy number and looked at when the checkpoint is reviewed.

What the rule does not do is make the model right more often. Temperature scaling changes no prediction; it changes how loudly the model says each one. A team that wants better accuracy still has to do the work of better data and better training, and a team that only refits the temperature has a model that is exactly as wrong as before and honest about it. That honesty is the entire point for anything the user reads as a claim.

Why a confident citation is worse than a hedged one

The place I learned to care about this was a helpdesk assistant that answered with grounded citations, built on open models fine-tuned with parameter-efficient adapters. A citation is a claim of certainty: here is the passage, this is where the answer comes from. When the model behind it is overconfident, the citation attaches that overconfidence to a specific document, and the user reads a wrong answer with a footnote. That is worse than a wrong answer with a hedge, because the hedge invites checking and the footnote discourages it. An assistant that says "I think it is this, see section four" when it is right 70 percent of the time is honest; one that says "It is this, see section four" with the same accuracy is not, and the difference between them is a temperature.

The fine-tune made the assistant better at the task and, in the same run, worse at knowing when it was wrong, and the second effect was invisible in the accuracy number that justified shipping. The reliability diagram was where it showed. Refitting the temperature after the final adapter merge brought the reported confidence back to something the citation could honestly carry, and the rule that it be refitted after every change was what kept it there through the next several releases, each of which changed the weights and would otherwise have quietly undone it.

The shipping checklist, in one line

Accuracy beside a reliability diagram, both computed after the last weight change, or the checkpoint does not ship. Fine-tuning will keep teaching models to sound sure, because that is what the loss rewards once the answers are settled. The temperature is how you take the sureness back down to what the model has earned, and "last" is the only place in the pipeline where it stays true.

CalibrationFine-TuningProbability
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS