Every wrong answer has an exchange rate
Accuracy rewards guessing. The fix is a penalty for wrong answers, and the penalty is not a research constant. It is a business number, and it sets the abstention threshold.
A helpdesk assistant that answers every question scores higher on accuracy than one that sometimes says it does not know. That is not a quirk of one benchmark. It is arithmetic. If a wrong answer and a declined answer both score zero, then answering is never worse than declining, and a model trained and selected on that scoreboard learns to guess. Kalai, Nachum, Vempala and Zhang put this at the centre of Why Language Models Hallucinate in September 2025: under a binary right-or-wrong scheme, guessing when unsure maximises the expected score, so the evaluations themselves keep hallucination alive.
Their proposed fix is to penalise confident errors more than abstentions. I agree with the fix and I want to push on the part the paper leaves open, because it is the part an engineer building a real system has to decide. How big should the penalty be. The answer is not a constant. It is an exchange rate between wrong answers and declined ones, it differs by deployment, and once you know it the abstention threshold follows from it.
Two models, three columns
The example that made the argument concrete came from OpenAI's own write-up of the paper, which compared two of its models on SimpleQA, a short-answer factuality test where each answer is graded correct, incorrect or not attempted, as described in the SimpleQA paper. One model abstained on 52% of questions, answered 22% correctly and got 26% wrong. The other abstained on 1%, answered 24% correctly and got 75% wrong. By accuracy alone the second model is ahead, 24 to 22. By any measure that counts a wrong answer as worse than silence, it is far behind.
The chart is the whole case for a third column. But look at what it does not tell you: which model to deploy. That depends on what a wrong answer costs relative to a declined one, and the chart has no opinion about that, because it cannot.
The exchange rate
Here is the number I keep next to any evaluation of an assistant that can decline. The wrong-answer exchange rate is the number of declined answers that one wrong answer is worth, in the deployment where the model will actually run. Call it R. It converts the three-column scorecard into a single cost: each declined answer costs 1, each wrong answer costs R, each correct answer costs 0, and the model with the lowest total is the one to ship.
R is not a property of the model. It is a property of what happens next. For a helpdesk with a human fallback, a decline costs a ticket: the person gets routed to someone who can answer. A wrong answer costs the same ticket, plus a repair conversation in which the person explains that they followed the assistant's instructions and it made things worse, plus some amount of trust that does not come back.
That is an R of several, and in a domain where following a wrong procedure can lock an account or corrupt a form, it is an R of many. For a trivia game with no consequences, R is close to 1, and the always-answer model is the right one. Same models, same chart, opposite decision.
Apply the rate to the two models. At R equal to 1, the abstaining model costs 26 plus 52, which is 78 per hundred questions, and the always-answer model costs 75 plus 1, which is 76. The guesser wins, narrowly. At R equal to 5, the abstaining model costs 130 plus 52, which is 182, and the guesser costs 375 plus 1, which is 376. The guesser loses by a factor of two. Neither result is visible in the accuracy column.
The threshold falls out of the rate
The exchange rate does more than rank models. If a model can report a confidence and abstain below a threshold, the rate tells you where to put the threshold, and the argument is short. Suppose the model's confidence is calibrated, so that an answer given with confidence c is right with probability c. Answering costs an expected R times (1 minus c). Declining costs 1. Answer whenever the first is smaller, which is whenever c exceeds 1 minus 1 over R. At R equal to 1 the threshold is zero and the model should always answer. At R equal to 5 it should answer only above 80% confidence. At R equal to 10, only above 90%.
The curve is a model and its assumption is strong: calibration. Real models are overconfident in places, and a fine-tuned model is often less calibrated than the base it came from. That does not break the rule; it changes the input. Measure the model's actual reliability curve on held-out questions, replace the straight line with it, and the threshold moves to wherever the measured curve says the expected wrong cost crosses 1. The exchange rate stays what it was, because it never depended on the model.
Where the rate comes from
When I was building the helpdesk assistant at CRIS, the thing that made abstention cheap was structural: a declined question went to a human queue that already existed. The assistant was in front of a process, not instead of one. That is the situation in which R is large and abstention is a feature, and it is the situation most internal assistants are in. The situation in which R is small is rarer than people assume: a consumer product with no fallback, where a shrug is as bad as an error because the user simply leaves.
Estimating R does not need a spreadsheet of costs, though one helps. It needs an honest answer to two questions. What happens after a decline, and what happens after a confident wrong answer. If the first is a ticket and the second is a ticket plus an apology plus a corrected form, R is at least 3. If the second can cause an action that someone has to undo, R is 10 or more, and the assistant should be quiet most of the time and precise when it speaks.
R also changes over the life of a deployment, and it is worth re-estimating when the surroundings change. Add a confirmation step before any action the assistant proposes, and wrong answers get cheaper, because a person now catches them; R falls and the threshold can drop. Remove the human queue behind the assistant to save money, and declines get more expensive, because a decline is now a dead end; R falls again, for the opposite reason. Both changes are made by people who never look at the model, and both should move the threshold. A rate written down once and never revisited is a threshold tuned for a system that no longer exists.
What this changes in the evaluation
Three things, and they are cheap. Report three columns, never one; an accuracy number with no abstention rate beside it is a number that rewards guessing, and you will select a guesser. Write R down for each deployment before you compare models, so that the comparison is a cost and not a taste. And set the abstention threshold from R and the measured reliability curve, rather than from a round number someone found in a demo.
The paper's authors argue that the field's scoreboards should change. Until they do, the scoreboard that matters is yours, and it needs one more number than it has.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS