2 September 2026 · 5 min read

Every paper has one load-bearing number

A paper's claim rests on one number that, a fifth worse, would have sunk it. Find that number first, ask four questions about how it was measured, and read the rest as context.

Every research paper I have read closely, and the one I have written, rests on a single number. It is the entry in one table that, had it come out a fifth worse, would have ended the submission, and everything else in the paper, the related work, the notation, the third ablation, exists to make that number credible. Reading a paper is the act of finding that number and checking how it was measured. Done in that order, it takes a quarter of the time the front-to-back approach does, and it catches the failures that the front-to-back approach reads straight past.

Why the number matters more than the prose

Kapoor and Narayanan's survey of leakage in machine-learning-based science found that across seventeen fields, 329 papers had a data leakage problem that undermined their central result, and the living version of the survey they maintain had grown to 648 papers across thirty fields by September 2026. In every one of those papers the prose was fine. The argument was coherent, the method was described, the related work was cited. The failure was in how one number was produced: a test set that overlapped the training set, a feature that would not exist at prediction time, a split that leaked the future into the past. The number said the method worked, and the number was wrong.

Papers with a documented leakage problem, by field Horizontal bars: law 156, molecular biology 67, medicine 57, radiology 55, neuropsychiatry 54, clinical epidemiology 48, neuroimaging 33, genomics 23, computer security 22, and 133 across the remaining twenty-one fields, for a total of 648 papers in thirty fields. Where the load-bearing number failed Papers with a leakage problem in the central result, by field, September 2026 Law 156 Molecular biology 67 Medicine 57 Radiology 55 Neuropsychiatry 54 Clinical epidemiology 48 Neuroimaging 33 Genomics 23 Computer security 22 21 other fields 133 648 papers in 30 fields. In each, the prose held and the number did not.
Source: the living leakage survey maintained by Kapoor and Narayanan, read in September 2026; the original 2022 paper counted 329 papers in 17 fields.

That is the case for reading numbers first. A paper's prose is written by people who believe the result, and it is persuasive in proportion to their belief. The number was produced by a procedure, and the procedure can be checked.

Finding it

The load-bearing number is rarely labelled. It is the cell in the results table that the abstract's claim points at, and the way to find it is to read the abstract, write down the claim in one sentence, and then go straight to the table that would have to contain the evidence for that sentence. Skip everything in between. The introduction, the related work and the method description are where the paper explains why the number should be believed; they are useless until you know which number.

The reading order, as a flow Five boxes in sequence: read the abstract and write the claim in one sentence; find the table cell the claim points at; read only the method section that produced that cell; ask the four questions; then, and only if the answers hold, read the rest as context. Number first, prose last Abstract claim, one line The cell the number Its method only what made it Four questions pass or bookmark The rest as context Related work, notation and ablations are read last, and only if the number holds. a quarter of the time, and it finds the failures front to back misses
Illustrative: the order I read in; it is the reverse of the order papers are written in.

The "fifth worse" test is how to be sure you have the right cell. Ask, for each candidate number: if this were twenty percent worse, would the paper still have been submitted? For most numbers in a results table the answer is yes; they are supporting evidence. For one, usually the comparison against the strongest baseline on the headline benchmark, the answer is no. That is the load-bearing number.

The four questions

Once found, the number gets four questions, and the answers decide what the paper is.

The four-question card beside a mock results table On the left, a small results table with three rows, a baseline at 71.2, a second baseline at 72.5, and the paper's method at 74.8, with the 74.8 cell highlighted as the load-bearing number. On the right, a card with four questions: what was compared, on how many items, with what variance, and who ran the baseline. One cell, four questions METHOD ACCURACY Baseline A, as reported 71.2 Baseline B, re-run 72.5 This paper 74.8 the cell a fifth worse would sink 1. What was compared, tuned how? 2. On how many items? 3. With what variance? Seeds, intervals. 4. Who ran the baseline? Four answers: a result to build on. Fewer: a preprint to bookmark. the table is a mock; the questions are the ones a reviewer asked of mine
Illustrative: a mock table and the four-question card; the values are invented to show the shape.

What was compared, and how was the comparison tuned? A method that beats a baseline the authors configured carelessly has beaten nothing. On how many items? A two-point gain on a thousand test items is a different claim from a two-point gain on fifty. With what variance? If the number is a single run, the gain may be a lucky seed, and a paper that reports no spread is asking to be trusted rather than checked.

And who ran the baseline? A baseline number copied from another paper, measured on a different split, with a different tokeniser or preprocessing, is not a comparison; it is two numbers that happen to share a column.

A paper that answers all four is a result to build on. A paper that cannot is a preprint to bookmark, which is not an insult. Most interesting ideas are first published before their number can be trusted, and the bookmark is the correct response: come back when someone has re-run it.

How a number goes wrong

The survey's own worked example shows how the load-bearing number fails while the prose holds. In civil war prediction, a series of papers reported that complex machine learning models substantially outperformed the logistic regression models political scientists had used for decades. The claim was specific, the tables were clear, and the argument in the text was reasonable. When Kapoor and Narayanan re-ran the comparisons, every one of the claimed improvements disappeared, because the pipelines had leaked information across the train-test split: imputation done on the whole dataset before splitting, or features that encoded the outcome. The load-bearing number in each paper was the gap between the complex model and the simple one, and the gap was an artefact of the procedure that produced it. No amount of reading the introduction would have found that. Reading the method section that produced the number, with the second question in hand, would.

It is also worth saying where the number lives when the paper has no results table. In a systems paper it is a latency or a throughput, usually in a figure rather than a table, and the four questions become: against what configuration, on what workload, over how many runs, and who tuned the baseline system. In a theory paper the analogue is the lemma the main theorem rests on, the one step in the proof that would collapse the result if it failed, and the questions become whether the lemma's conditions match the theorem's, and whether the step is proved or cited. The shape is the same: one place carries the weight, and the reading starts there.

What carried my own paper

I say this as someone who has been on the other side of the questions. My paper on CYK parsing at the ACL 2026 Student Research Workshop made a claim that rested on one table: the benchmark comparison between CYK and three other parsers, LR(1), Earley and recursive descent, on a shared grammar. Every other section of the paper, the background on normal forms, the complexity analysis, the discussion, was context for that table. A reviewer who read it the way this post describes would have gone to the table first, and what the table needed in order to be trusted was exactly the four answers: the same grammar and sentences for every parser, the number of sentences, the spread across runs, and the fact that I had run all four parsers myself rather than quoting three of the numbers from elsewhere.

Knowing that the table was the paper changed how I wrote it. The method section that produced the table got more care than the introduction, the variance got reported even where it was boring, and the sentence in the abstract was written to point at the cell. It is the discipline the classic three-pass reading method gets at from the reader's side, and the load-bearing number is what it is looking for. Find it first. Everything else in the paper is there to help you check it.

ResearchReading PapersEvaluation
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS