Crawler Access

Check that AI crawlers can reach your site and read what is on it.

Crawler Access in Proofsource checks the two things that decide whether AI can learn from your site: whether GPTBot, ClaudeBot, PerplexityBot and the other AI crawlers are allowed in and actually served, and whether the page they receive contains your words or an empty JavaScript shell.

2 free scans, today and tomorrow. No card · no sales call

robots.txt

A verdict for every AI crawler in your robots.txt.

Proofsource reads your robots.txt and gives a verdict for ten AI crawlers. Knowing which crawler does what matters: OAI-SearchBot and PerplexityBot fetch pages for live answers, while GPTBot and Google-Extended govern training data, so you can block training without disappearing from AI search.

  • GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, Applebot-Extended, CCBot and Bytespider
  • Allowed or disallowed, and not determined when the file cannot be read, never a guess
  • Google-Extended is a robots.txt token, not a separate crawler, and is checked as one
Server check

robots.txt is a request. The server decides who gets in.

A site can allow GPTBot in robots.txt and still answer it with a 403 from a CDN bot firewall. Proofsource requests your homepage as GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot and PerplexityBot and records what your server actually returned. When we ran it on our own site, two of the five were turned away while robots.txt allowed all of them.

  • Served, turned away or unclear, per crawler
  • The status code each crawler received
  • Catches CDN and firewall blocks that robots.txt checks miss
JavaScript rendering

See how much of your page a bot can actually read.

Many AI crawlers read the raw HTML and do not run JavaScript. Proofsource compares the words a browser renders with the words in the raw HTML. On 4 Oct 2026 our own homepage showed 2,416 words in a browser and 9 to a bot. If your content only appears after JavaScript runs, most AI crawlers never see it.

  • Rendered words against raw-HTML words, per page
  • Flags pages with no body text, no H1 or no crawlable links
  • Re-check after you ship server rendering or prerendering
Why it matters

If AI cannot read you, it cannot recommend you.

Access

Find silent blocks

CDN bot settings can turn AI crawlers away without anyone on your team noticing. The server check shows it.

Content

Make your words visible

A page that renders only in JavaScript is close to blank for most AI crawlers. The rendering check measures how close.

Control

Block training, keep search

A verdict per crawler shows exactly which AI systems you let in, so blocking training is a choice, not an accident.

Questions

Questions buyers ask

How do I check whether GPTBot can access my website?

Check two things. First, robots.txt: it should not disallow GPTBot. Second, the server: request a page with GPTBot's user agent and confirm you get a 200, not a 403 from a CDN or firewall. Proofsource runs both checks and also confirms the page contains readable text.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects content that may be used to train OpenAI models. OAI-SearchBot fetches pages for ChatGPT search results, and ChatGPT-User fetches a page when a user's conversation needs it. Blocking GPTBot does not remove you from ChatGPT search, but blocking OAI-SearchBot does.

What is Google-Extended?

Google-Extended is a robots.txt token, not a separate crawler. Disallowing it tells Google not to use your content for Gemini and Vertex AI model training. It does not affect Google Search.

Can AI crawlers read JavaScript websites?

Many cannot. Several AI crawlers read only the raw HTML a server returns and do not run JavaScript, so a client-rendered page can look almost empty to them. Server-side rendering or prerendering puts the text in the HTML where every crawler can read it.

Why would my site block AI crawlers without my knowing?

CDN and security providers offer bot-blocking settings, and some block AI training crawlers by default or with one switch. The block happens at the server, so robots.txt still says the crawler is allowed. Only a request made as the crawler shows it.

Can I try Proofsource for free?

Yes. The free trial is 2 free scans, today and tomorrow: 25 buyer questions on ChatGPT, Perplexity, Claude and Google AI Overviews. No card. After the second scan, scanning pauses and your answers stay in your account.

How much does Proofsource cost?

Custom plans start from $64 per brand per month. Final price confirmed at launch. Every plan includes the four core engines; you choose how many questions and brands to track.

Features

More of the platform

Know

Ask Proof

An assistant that answers questions about your AI visibility from your own data, and shows the queries behind every number.Explore Ask Proof

Know

AI Visibility

Mention rate, share of voice and average position on each AI engine, with a 95% Wilson interval on every number.Explore AI Visibility

Know

Prompts

The buyer questions you track, grouped by topic, tagged branded or unbranded, ranked by search demand, with the searches each engine runs.Explore Prompts

See every feature

Make sure AI can read you.

See whether AI names your brand, who it names instead, and what to fix. 2 free scans, today and tomorrow. No card · no sales call.

Last updated