Write the page for two readers
The 2026 evidence on llms.txt is in, and it says the file is barely read. Every crawler that matters reads the HTML and the structured data, so the work belongs in the page itself.
This site has an llms.txt. So does the marketing site I built for a legal practice. I put them there in the spirit of the proposal: a small markdown file that tells a language model what a site is about, in the way robots.txt tells a crawler where it may go. In 2026 the evidence on whether anything reads those files arrived, and it is worth being honest about what it says, because the conclusion is not that the effort was wasted. It is that the effort was in the wrong file.
What the logs say
Ahrefs analysed server logs across 137,210 domains in May 2026 and found that 97 percent of llms.txt files received no requests at all. Of the requests that did arrive, most came from bots that are not AI tools: SEO auditors, monitoring services, and the study's own crawlers. GPTBot accounted for about 4.5 percent of llms.txt fetches and ClaudeBot under one percent, with AI retrieval bots in aggregate near one percent. The file that was supposed to be the AI reader's front door was, on the evidence, mostly visited by tools checking whether it existed.
Google said the same thing from the other side. In June 2026 its Search Central documentation added a clarification that sites do not need new machine-readable files, AI text files or markdown to appear in Google Search, including its generative features, because Google Search does not use them. The file is fine to have. It does nothing there.
What the crawlers actually read
The reason is not that language-model companies are ignoring site owners. It is that they never needed the file. A crawler that indexes pages for retrieval reads HTML, because HTML is where the content is, and reads structured data, because JSON-LD is where the facts are in a form a machine can trust. Those two things existed before llms.txt and they are what every reader that matters, search engines, answer engines, retrieval crawlers, was already consuming. A separate markdown summary of the site is a second copy of information the page already carries, maintained by hand, drifting from the page from the day it is written.
That is the framing I now use for every page on a site I build, and I call it the two-reader page. Every page is written once for a person and once, in the same document, for a machine. The prose carries the argument for the person. The JSON-LD carries the facts for the machine: who, what, when, which organisation, which page this is a part of. Clean semantic HTML lets a crawler tell the article from the navigation. llms.txt, if it exists at all, is a courtesy index for the second reader, and nothing goes in it that is not already on a page.
What the second reader needs from the page
Writing for the machine is more concrete than it sounds, and most of it is discipline rather than markup. The facts a machine reader wants are the ones it cannot infer from prose with confidence: who wrote this, when, on behalf of which organisation, what the page is about in one line, and how it relates to the other pages on the site. JSON-LD carries all of that in a form that does not depend on parsing English. On this site the home page carries a single graph in which the person, the company, the services brand, the website, the research paper and the page itself are separate nodes that reference one another by identifier, so that a reader arriving at any node can walk to the rest.
The prose then has to agree with the graph, which is the discipline part. If the structured data says the paper was selected for a workshop and the prose says it was published in a journal, the machine reader has two facts and no way to choose; the person reading the prose has one fact and no idea it is contested. Keeping both in one document, edited together, is what stops that drift, and it is the reason the side-file model fails: the file is edited on a different day, by a different reflex, and nobody reads it back against the page.
Semantic HTML is the third piece and the cheapest. An article element around the article, a nav element around the navigation, headings in order, a time element with a machine-readable date. None of it is new and all of it is skipped in a rush, and a crawler faced with a page of divs falls back to guessing which text is the content, which is the failure that makes a site's own summary look attractive in the first place.
The one file-level signal with teeth
There is one place where a file-level line does something, and it is the oldest one. In July 2026 Cloudflare announced that it would require AI companies to separate the crawlers they use for search from those used for training and for agents, with a deadline of 15 September, after which training and agent crawlers would be blocked by default on many sites while search crawlers remained allowed; its press release frames it as ending the trade-off between being found and being trained on. The mechanism that expresses a site's choice is robots.txt and the network's own controls, not a markdown summary. The signal that has consequences is the one that says who may fetch, and it lives in the file crawlers have honoured for thirty years.
What I changed on this site
Two things, and one thing I deliberately did not. The JSON-LD graph on the home page became the source of truth for the facts about me, the company and the work: one graph, with the person, the organisation, the site, the pages and the publication linked to one another by identifier, so that a machine reading any page can walk to the others. The article pages carry their own structured data with dates, authorship and the parent blog, and the HTML around the prose is semantic enough that the article is distinguishable from the navigation without heuristics.
The llms.txt stayed, rewritten as an index of pages that already exist, with nothing in it that the pages do not say. It costs nothing to keep, a few readers do fetch it, and the discipline of not putting anything new in it is what stops it drifting. It is a table of contents, not a second site.
And the robots.txt line got the attention it had been missing, because it is the one file-level statement with consequences: which crawlers may read, for which purposes, is a decision, and after September it is a decision with defaults set by someone else if the site does not set its own.
The rule
Write every page for two readers, in one document. Put the argument in the prose and the facts in the structured data, and keep them consistent because they are on the same page. Treat any side file as a courtesy that must never contain something a page does not, and treat robots.txt as the only file that governs. The machine reader was never waiting at a special door. It came in the front, the same as everyone, and the page is what it read.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS