Is your website blocking ChatGPT without you knowing?
It might be. Since July 2025 Cloudflare, which sits in front of a large share of websites, asks new sites whether to allow AI crawlers and blocks them unless the owner says yes. Some security plugins and robots.txt templates do the same. If the crawlers behind ChatGPT, Claude or Perplexity are blocked, those assistants can't read your pages and describe you from other sources instead.
Two kinds of AI crawlers
This is the part most articles skip. AI companies run different crawlers for different jobs, and you can treat them differently:
| Crawler | Company | What it does |
|---|---|---|
| OAI-SearchBot | OpenAI | Finds pages to show in ChatGPT search answers |
| ChatGPT-User | OpenAI | Fetches a page when a ChatGPT user asks about it |
| GPTBot | OpenAI | Collects pages to train future models |
| ClaudeBot | Anthropic | Collects pages for Claude |
| PerplexityBot | Perplexity | Indexes pages for Perplexity answers |
| Google-Extended | A switch for Gemini model training. Blocking it does not remove you from Google Search | |
| Applebot-Extended | Apple | A switch for Apple's AI training. Regular Applebot still powers Siri and Spotlight |
For a local business the logic is simple: search and answer crawlers bring customers, so let them in. Training crawlers are a matter of taste. Blocking all of them together is the mistake.
How sites end up blocking AI without meaning to
- Cloudflare's default. On July 1, 2025 Cloudflare announced it would block AI crawlers by default and ask every new domain whether to allow them. An owner, or the person who set up the site, clicks through setup and the door is closed.
- Security plugins and "bot protection". Firewalls often treat any unfamiliar crawler as a threat.
- A copied robots.txt. Lists of "bad bots" circulating online often include AI crawlers alongside actual scrapers.
- Nobody looked. The block produces no error and no email. The site works perfectly for people. Only the assistants go quiet.
How to check
- Open
yourdomain.com/robots.txtand look for the names above next toDisallow: /. - If your site uses Cloudflare, look in the dashboard under Security for the AI crawlers setting.
- Ask ChatGPT with search on to summarise a specific page of your site. If it says it can't access it while the page loads fine for you, something is blocking it.
A robots.txt that welcomes search and answer crawlers and says no to training looks like this:
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: *
Allow: /
Sitemap: https://yourdomain.com/sitemap.xml
Robots.txt is a request, not a lock. Firewall rules override it, so check both.
Check this on your own business
The free 60-second Check-Up reads your website and Google profile and shows where you stand on this: robots.txt, Quotable passages. Public data only, no login, no card.
Run the free Check-Up →Or we run it for you
We set your robots.txt and hosting rules so AI search crawlers can read you, keep training crawlers your choice, and check it every week so a plugin or hosting update doesn't quietly close the door again. Month to month, a new website included, plans from $199 a month.
See how it works →Questions
Will blocking GPTBot remove me from ChatGPT?
Not by itself. GPTBot is for training. ChatGPT search uses OAI-SearchBot and ChatGPT-User. Block those and you disappear from ChatGPT's live answers.
Does blocking Google-Extended hurt my Google ranking?
No. Google says Google-Extended only controls use in Gemini models and has no effect on Google Search.
Should a local business allow AI training crawlers?
It's your choice. There is no proven ranking benefit either way. Allowing search and answer crawlers is the part that matters for being recommended.