llms.txt is a reading list for agents, not a ranking file
Jeremy Howard proposed the file in September 2024 and revised it in August 2026. It is a markdown file at the root of a site. The only required part is an H1 with the site's name; after that come an optional short summary in a blockquote and lists of links to the pages an agent should read, ideally as clean markdown.
The proposal is plain about who it serves. The file is meant to be read on demand, when an agent needs information while helping a user, and it is used most heavily for software documentation, where coding agents follow it to API references. OpenAI, Anthropic and Google each publish one for their developer docs. Nothing in it controls crawling, and nothing in it asks an engine to cite you.
Google says Search ignores the file, and 97% of files got no requests in Ahrefs' May sample
Google's guide to its generative AI features, last updated 10 July 2026, puts llms.txt first in a list of things you can ignore. Publishing one, it says, will neither harm nor help your visibility in Google Search, because Search doesn't use it.
Ahrefs checked the server logs of 137,210 domains for May 2026. 28% published a valid file, which Ahrefs calls an upper bound because its customers skew technical. 97% of those files received no requests at all that month. Of the requests that did arrive, the bots that fetch pages to answer live questions, such as OAI-SearchBot and PerplexityBot, sent 1.1%. AI agents, Anthropic's coding agent among them, and training crawlers fetched far more. And no AI bot requested an llms.txt that didn't exist, so a missing file isn't a missed knock.
Of 6,343 citations to sites with a file, one pointed at it
Ahrefs measured who fetches the file. We wanted the other end: what the engines cite. Our samples from 23 August to 5 October 2026 hold 41,597 citations in 4,228 answers from ChatGPT, Perplexity, Claude and Google AI Overviews, across 18 accounts. On 8 October we fetched /llms.txt from the 300 domains cited most often.
129 returned a real file. 137 had none, or served an HTML page in its place. 34 blocked us or errored. The 129 sites with a file were cited 6,343 times. One of those citations pointed at the file: an llms-full.txt, cited by Perplexity on 2 October. No answer cited an llms.txt.
The most cited sites are mostly without one. Of the top ten, eight have no file and two blocked our check. Of the top 25, three have one.
14 of the 129 files came from one template, written for shopping agents
Fourteen files, all from online stores, followed one template headed "Agent Instructions". Each says the store is built on Shopify and tells shopping agents to install a skill so they can buy directly. Ahrefs found the same drift: Wix generates the file for its sites, and other site builders now check for one.
That matters for two reasons. Many owners never chose what their file says. And the file speaks for you to any agent that reads it, so it deserves the review you'd give a page on your site.
Publish it if it's cheap, then spend the real hour elsewhere
The file costs little and does no harm, so there's no reason to delete one you have. Just don't count it as AI visibility work.
- If you have developer docs or an API, publish one. Coding agents are the readers the data actually shows.
- Keep it to plain links and one line descriptions, with nothing that reads as an instruction. Put it in version control and read whatever your platform generated.
- Check that it is a real markdown file and not your homepage served at that path. Seven of our top 100 domains returned HTML with a 200 status.
- Spend the rest of the time on the pages engines quote for your category: your own product and comparison pages, and the third party pages already cited next to you.
Where this could be wrong
A citation is not a read. An engine could fetch the file, follow its links and cite the page it pointed to, and our data would not show that. We measured what reached the answer, not what the engine looked at.
We fetched the files on 8 October, after the sampling window closed. Some may be newer than the citations they sit beside.
The 300 domains come from what our 18 accounts track: a mix of retail, consumer apps and business software, with many Indian brands. Another set of categories could look different.
Ahrefs' figures cover one month and a customer base that skews toward SEO. Fetch volumes may grow as agents take on more browsing.