8 March 2026 · 5 min read

Every wrong answer has an exchange rate

Accuracy rewards guessing. The fix is a penalty for wrong answers, and the penalty is not a research constant. It is a business number, and it sets the abstention threshold.

A helpdesk assistant that answers every question scores higher on accuracy than one that sometimes says it does not know. That is not a quirk of one benchmark. It is arithmetic. If a wrong answer and a declined answer both score zero, then answering is never worse than declining, and a model trained and selected on that scoreboard learns to guess. Kalai, Nachum, Vempala and Zhang put this at the centre of Why Language Models Hallucinate in September 2025: under a binary right-or-wrong scheme, guessing when unsure maximises the expected score, so the evaluations themselves keep hallucination alive.

Their proposed fix is to penalise confident errors more than abstentions. I agree with the fix and I want to push on the part the paper leaves open, because it is the part an engineer building a real system has to decide. How big should the penalty be. The answer is not a constant. It is an exchange rate between wrong answers and declined ones, it differs by deployment, and once you know it the abstention threshold follows from it.

Two models, three columns

The example that made the argument concrete came from OpenAI's own write-up of the paper, which compared two of its models on SimpleQA, a short-answer factuality test where each answer is graded correct, incorrect or not attempted, as described in the SimpleQA paper. One model abstained on 52% of questions, answered 22% correctly and got 26% wrong. The other abstained on 1%, answered 24% correctly and got 75% wrong. By accuracy alone the second model is ahead, 24 to 22. By any measure that counts a wrong answer as worse than silence, it is far behind.

Two models on SimpleQA, three outcomes each Two stacked bars. The first model: 22% correct, 26% wrong, 52% abstained. The second model: 24% correct, 75% wrong, 1% abstained. The second model has slightly higher accuracy and nearly three times the error rate. Accuracy hides the third column Share of SimpleQA questions by outcome Model that abstains when unsure 22% correct 26% wrong 52% abstained Model that always answers 24% correct 75% wrong correct wrong abstained (1% on the second bar)
Source: the model comparison in OpenAI's summary of Kalai et al. (2025), on the SimpleQA evaluation.

The chart is the whole case for a third column. But look at what it does not tell you: which model to deploy. That depends on what a wrong answer costs relative to a declined one, and the chart has no opinion about that, because it cannot.

The exchange rate

Here is the number I keep next to any evaluation of an assistant that can decline. The wrong-answer exchange rate is the number of declined answers that one wrong answer is worth, in the deployment where the model will actually run. Call it R. It converts the three-column scorecard into a single cost: each declined answer costs 1, each wrong answer costs R, each correct answer costs 0, and the model with the lowest total is the one to ship.

R is not a property of the model. It is a property of what happens next. For a helpdesk with a human fallback, a decline costs a ticket: the person gets routed to someone who can answer. A wrong answer costs the same ticket, plus a repair conversation in which the person explains that they followed the assistant's instructions and it made things worse, plus some amount of trust that does not come back.

That is an R of several, and in a domain where following a wrong procedure can lock an account or corrupt a form, it is an R of many. For a trivia game with no consequences, R is close to 1, and the always-answer model is the right one. Same models, same chart, opposite decision.

Apply the rate to the two models. At R equal to 1, the abstaining model costs 26 plus 52, which is 78 per hundred questions, and the always-answer model costs 75 plus 1, which is 76. The guesser wins, narrowly. At R equal to 5, the abstaining model costs 130 plus 52, which is 182, and the guesser costs 375 plus 1, which is 376. The guesser loses by a factor of two. Neither result is visible in the accuracy column.

The same scorecard under two exchange rates A table with two models as rows and columns for correct, wrong and declined shares, followed by two cost columns. At an exchange rate of 1, the abstaining model costs 78 and the always-answer model 76, so the guesser wins. At an exchange rate of 5, the abstaining model costs 182 and the guesser 376, so the abstaining model wins. Cost per 100 questions = wrong times R + declined MODEL CORRECT WRONG DECLINED R = 1 R = 5 Abstains when unsure 22 26 52 78 182 Always answers 24 75 1 76 376 The dot marks the cheaper model: the guesser at R = 1, the abstainer at R = 5. Accuracy, the first column, ranks them the same way under both rates. shares from the chart above; costs are arithmetic on those shares
Illustrative: the costs are computed from the published shares under two stated exchange rates; the shares are from OpenAI's comparison.

The threshold falls out of the rate

The exchange rate does more than rank models. If a model can report a confidence and abstain below a threshold, the rate tells you where to put the threshold, and the argument is short. Suppose the model's confidence is calibrated, so that an answer given with confidence c is right with probability c. Answering costs an expected R times (1 minus c). Declining costs 1. Answer whenever the first is smaller, which is whenever c exceeds 1 minus 1 over R. At R equal to 1 the threshold is zero and the model should always answer. At R equal to 5 it should answer only above 80% confidence. At R equal to 10, only above 90%.

Expected cost against abstention threshold under two exchange rates Two curves over thresholds from 0 to 1. At an exchange rate of 1 the cost per 100 questions rises from 50 at threshold zero to 100 at threshold one, so the best threshold is zero. At an exchange rate of 5 the cost falls from 250 at threshold zero to a minimum of 90 at threshold 0.8 and rises to 100 at threshold one. Where to put the abstention threshold Cost per 100 questions, in declined-answer units, calibrated model R = 5 R = 1 250 200 150 100 50 0 0 0.2 0.4 0.6 0.8 1.0 Confidence below which the model declines minimum at 0.8 minimum at 0
Illustrative: computed from a calibrated model whose per-question confidence is spread evenly from 0 to 1, so cost per question is the declined share plus R times the expected wrong share; the construction is stated in the text.

The curve is a model and its assumption is strong: calibration. Real models are overconfident in places, and a fine-tuned model is often less calibrated than the base it came from. That does not break the rule; it changes the input. Measure the model's actual reliability curve on held-out questions, replace the straight line with it, and the threshold moves to wherever the measured curve says the expected wrong cost crosses 1. The exchange rate stays what it was, because it never depended on the model.

Where the rate comes from

When I was building the helpdesk assistant at CRIS, the thing that made abstention cheap was structural: a declined question went to a human queue that already existed. The assistant was in front of a process, not instead of one. That is the situation in which R is large and abstention is a feature, and it is the situation most internal assistants are in. The situation in which R is small is rarer than people assume: a consumer product with no fallback, where a shrug is as bad as an error because the user simply leaves.

Estimating R does not need a spreadsheet of costs, though one helps. It needs an honest answer to two questions. What happens after a decline, and what happens after a confident wrong answer. If the first is a ticket and the second is a ticket plus an apology plus a corrected form, R is at least 3. If the second can cause an action that someone has to undo, R is 10 or more, and the assistant should be quiet most of the time and precise when it speaks.

R also changes over the life of a deployment, and it is worth re-estimating when the surroundings change. Add a confirmation step before any action the assistant proposes, and wrong answers get cheaper, because a person now catches them; R falls and the threshold can drop. Remove the human queue behind the assistant to save money, and declines get more expensive, because a decline is now a dead end; R falls again, for the opposite reason. Both changes are made by people who never look at the model, and both should move the threshold. A rate written down once and never revisited is a threshold tuned for a system that no longer exists.

What this changes in the evaluation

Three things, and they are cheap. Report three columns, never one; an accuracy number with no abstention rate beside it is a number that rewards guessing, and you will select a guesser. Write R down for each deployment before you compare models, so that the comparison is a cost and not a taste. And set the abstention threshold from R and the measured reliability curve, rather than from a round number someone found in a demo.

The paper's authors argue that the field's scoreboards should change. Until they do, the scoreboard that matters is yours, and it needs one more number than it has.

HallucinationEvaluationHelpdesk
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS