Method5 min read

Blocking GPTBot doesn't remove you from ChatGPT. Blocking OAI-SearchBot does.

Every engine now runs several crawlers with different jobs, and the one most people block is not the one that decides whether you show up. Here's the current map, from each company's own documentation, and the checks that tell you where you stand.

OpenAI sends four agents, and only one gates ChatGPT search

There's a robots.txt snippet going around that disallows GPTBot. Plenty of people pasted it in and felt they'd settled the ChatGPT question. They settled a question. Just not that one.

OpenAI documents four user agents, each with its own job. GPTBot crawls content that may be used to train the foundation models. OAI-SearchBot is the one used to surface sites in ChatGPT's search features, and OpenAI's own page says sites opted out of it will not be shown in ChatGPT search answers, though they can still turn up as navigational links. ChatGPT-User fetches a page when a person asks ChatGPT something, and because a user started it, OpenAI says robots.txt rules may not apply. OAI-AdsBot only visits pages submitted as ads.

OpenAI is explicit that each setting is independent. You can allow OAI-SearchBot and disallow GPTBot, which says yes to being findable and no to being training data. That is a reasonable position for most companies to hold, and hardly any robots.txt file we read actually says it. Give it about a day either way: OpenAI says it takes roughly 24 hours from a robots.txt change for their systems to catch up.

Anthropic splits it three ways, and Perplexity ignores you when a person asks

Anthropic runs three. ClaudeBot collects web content that may go into training. Claude-SearchBot indexes content to improve search results. Claude-User fetches a page when someone asks Claude a question. Anthropic's help page spells out the consequence of each: turn off Claude-SearchBot and your content stops being indexed for search, turn off Claude-User and Claude stops retrieving your page in response to a user's question.

Perplexity runs two. PerplexityBot exists to surface and link sites in Perplexity's results and is not used to collect training data at all. Perplexity-User visits a page when someone asks a question, and Perplexity's docs say the quiet part out loud: since a user requested the fetch, that fetcher generally ignores robots.txt.

Add it up and the blanket block-the-AI-bots rule lands somewhere strange. Of the nine agents above, most of the ones people mean to stop are the ones that put your link in front of a buyer, and one of them isn't listening to robots.txt anyway.

Google-Extended has nothing to do with AI Overviews

This is the one that costs people the most. Google-Extended is a robots.txt token, with no user agent string of its own, controlling whether crawled content trains future Gemini models and grounds Gemini Apps. Google's crawler documentation states flatly that it does not affect a site's inclusion in Google Search and is not used as a ranking signal. Blocking it does nothing to AI Overviews.

AI Overviews and AI Mode are Search. Google says so in its own guidance for site owners: AI is built into Search, and robots.txt directives for Googlebot are the control. Then comes the sentence that catches people. To be eligible as a supporting link in AI Overviews or AI Mode, a page has to be indexed and eligible to be shown with a snippet. If you or someone before you set nosnippet, data-nosnippet or a tight max-snippet to keep scrapers off, you removed yourself from the AI answer, and nothing in your analytics reported it.

The same Google guide says you don't need to create new machine readable files or AI text files to appear in these features. We've been telling people not to treat llms.txt as a fix for a while: none of the engines we watch documents one as a ranking signal. It's useful to have Google state the general version in its own words.

Your CDN may have answered this for you already

On 1 July 2025 Cloudflare announced it was changing the default to block AI crawlers unless they pay for content, and shipped a managed robots.txt that it serves on your behalf and keeps updated as new bots appear. Every new site onboarding to Cloudflare is now asked to make this choice at setup.

So for a lot of companies, the honest answer to what does my robots.txt say about ChatGPT is I don't know, my CDN writes it. A WAF rule does the same job more thoroughly. The request never reaches your origin, so nothing in your logs or your analytics records a visit that didn't happen. Your page still ranks, still loads, still looks right to you.

And a fetch that lands still has to come back with words in it. A client-rendered page hands these agents a 200 and an empty container, which reads as a working page to you and as nothing at all to them.

Six checks that settle it in an afternoon

None of this needs a tool. It needs curl, your logs, and an hour. We run it as a bot-accessibility audit with Cloudflare crawler analytics behind it, but you can do every line yourself.

  • Fetch your robots.txt over the public internet, not from your repo, and search it for each token by name: OAI-SearchBot, GPTBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Google-Extended, Googlebot.
  • Do it for every subdomain. docs.example.com and www.example.com are separate files and separate decisions, and the docs subdomain is usually the one you want cited.
  • Request a page you want quoted using each agent's published user agent string, and read what comes back. A 200 with an empty div is a failure, not a pass.
  • Look in your CDN and WAF for a managed robots.txt or a bot rule nobody on the current team wrote.
  • Check the head of your best pages for noindex, nosnippet, data-nosnippet and max-snippet. Any of those can quietly remove a page from AI Overviews eligibility.
  • Grep your access logs for the agent names, then check the source addresses against the IP ranges OpenAI, Perplexity and Anthropic publish. A user agent string is free to type. The address range is the part that verifies.

Where this could be wrong

Everything above is each engine describing its own crawlers, and both the tokens and the behaviour move. Anthropic's page carries an April 2026 date, Google's crawler list was last updated in July 2026, and OpenAI adds an agent whenever it ships a product that fetches pages. Read the live docs before you change a line, not this article.

robots.txt is a request. A company saying its bot honours it is a policy statement and not an enforcement mechanism, and two of the fetchers here document that they don't apply it to user-initiated requests at all. If you need something not read, robots.txt was never the tool.

Letting a crawler in also doesn't get you cited. Access is the floor, not the answer. In the category we sampled, listicles were 44% of what ChatGPT cited for a best-of question, so being readable and being chosen are two different problems. Our 44% is a single measurement on the sample we took, not a constant.

Still worth an afternoon, because this failure is silent by construction. Nothing anywhere reports a fetch that never happened.

Sources

Where the outside numbers come from

  1. Overview of OpenAI Crawlers · OpenAI
  2. Does Anthropic crawl data from the web, and how can site owners block the crawler? · Anthropic
  3. Perplexity Crawlers · Perplexity
  4. Google's common crawlers · Google Search Central
  5. AI features and your website · Google Search Central
  6. Content Independence Day: no AI crawl without compensation! · Cloudflare
All notes

Questions

Should I block GPTBot?

That's a business call about training data, not a visibility one. Disallowing GPTBot says don't train on me. It does not remove you from ChatGPT's answers, because OAI-SearchBot is the token that governs search. The two settings are independent, so you can hold both positions at once.

Is an llms.txt file worth publishing?

Not as a visibility fix. No engine we watch documents one as a ranking signal, so publish one only if it's cheap. Google says in its own guidance that you don't need new machine readable or AI text files to appear in AI Overviews or AI Mode.

How do I know a request claiming to be GPTBot really is?

Check the source address. OpenAI publishes IP ranges for each of its agents as JSON, and Perplexity and Anthropic publish theirs too. Anyone can send a user agent header that says GPTBot, so match on the address range before you allow or count anything.

Keep reading

Related notes

Your own numbers beat our best post.

The shortlist in your category is being written right now, whether anyone reads this page or not. A free trial, 25 questions on four engines today and tomorrow, tells you if your name is in it.