robots.txt checker for AI crawlers and Googlebot

A free robots.txt checker. Enter a domain and we read its robots.txt for 16 crawlers, the AI bots from OpenAI, Anthropic and Perplexity plus Googlebot and Bingbot, and show whether each one is allowed, blocked or not mentioned, with the line that decides it.

Sixteen crawlers, AI and search, and what each one is for.

Results show here. Nothing is stored, and no AI model is involved: it is a plain robots.txt read.

  • GPTBot · OpenAI · Training

    Collects pages to train OpenAI models.

  • OAI-SearchBot · OpenAI · Search

    Indexes pages so ChatGPT search can cite them.

  • ChatGPT-User · OpenAI · User fetch

    Opens a page when a ChatGPT user asks it to.

  • ClaudeBot · Anthropic · Training

    Collects pages to train Claude models.

  • Claude-SearchBot · Anthropic · Search

    Indexes pages so Claude can cite them in search answers.

  • Claude-User · Anthropic · User fetch

    Opens a page when a Claude user asks it to.

  • anthropic-ai · Anthropic · Training

    An older Anthropic token that many robots.txt files still name.

  • PerplexityBot · Perplexity · Search

    Indexes pages so Perplexity can cite them in answers.

  • Perplexity-User · Perplexity · User fetch

    Opens a page when a Perplexity user asks it to.

  • Google-Extended · Google · Training

    Decides whether Gemini may train on pages Googlebot crawls. Search is unaffected.

  • Googlebot · Google · Search

    Google Search, the source AI Overviews and AI Mode draw on.

  • Bingbot · Microsoft · Search

    Bing search, the index Copilot answers draw on.

  • Applebot-Extended · Apple · Training

    Decides whether Apple may train its models on pages Applebot crawls.

  • CCBot · Common Crawl · Training

    Builds the open Common Crawl dataset many models train on.

  • Meta-ExternalAgent · Meta · Training

    Collects pages to train Meta AI models.

  • Bytespider · ByteDance · Training

    Collects pages to train ByteDance models.

Rules follow RFC 9309: a crawler obeys the groups that name it, otherwise the * group; the longest matching rule wins, and Allow wins a tie. The verdict is for the homepage. Training crawlers feed future models; search crawlers decide whether an AI answer can cite you today; user fetches happen when someone asks an assistant to open your page.

robots.txt is only half of access. A CDN bot setting can still turn a crawler away that robots.txt allows, and this check cannot see that; Crawler access in Proofsource requests your pages as each crawler to find out. Read why blocking GPTBot does not remove you from ChatGPT, check your file with the llms.txt validator, or start with the guide to AI visibility.

What the answers mean.

Should I block training crawlers?

That is a business call. Blocking GPTBot or ClaudeBot keeps your pages out of future training sets, but it does not stop ChatGPT or Claude from citing you through their search crawlers. Blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot does remove you from those answers.

Why does a crawler say not mentioned?

Your robots.txt has no group for it, so it follows the * group. If there is no * group either, nothing restricts it.

Does robots.txt stop every bot?

No. It is a request that the major AI companies say they honour. It does not block a crawler that ignores it; a firewall rule does that.

Being crawlable is step one.

Proofsource shows whether ChatGPT, Perplexity, Claude and Google actually name you, which competitors they name instead, and what to fix.