GEO & AI10 min read

AI Crawlers (GPTBot, ClaudeBot, PerplexityBot…): Should You Block Them?

GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended: what each AI crawler does, what to block, and 3 annotated robots.txt setups.

By Valentin Halgand, founder of Agence Zen

AI Crawlers (GPTBot, ClaudeBot, PerplexityBot…): Should You Block Them?
In short. No, not all of them. Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and the agents that read a page when a user asks decide whether you appear in ChatGPT, Claude or Perplexity answers: block them and you drop out. Training crawlers (GPTBot, ClaudeBot, CCBot…) can be refused separately; OpenAI states that these settings are independent.

If your website is public, AI crawlers are almost certainly reading it already. Some sites have blocked all of them, sometimes unknowingly, through a hosting or CDN setting. For a business that wants ChatGPT, Claude, Perplexity, Gemini or Copilot to recommend it, that is an expensive mistake: an engine that cannot read your pages cannot cite them. The researchers who formalised GEO (Aggarwal et al., KDD 2024) measured visibility gains of up to 40% in generative engine responses, with results varying by domain. None of that works unless the engine can reach your content, which is what this guide covers. For the wider strategy, see our complete GEO guide.

Three Kinds of AI Crawler, Three Decisions

Vendors now document three distinct jobs. Mixing them up is how sites end up blocking what makes them visible to protect something that was never at stake.

  • Training (GPTBot, ClaudeBot, CCBot): collecting public pages to train future models. Refusing them keeps your future content out of training data. It does not keep your business out of answers.
  • Search indexing (OAI-SearchBot, Claude-SearchBot, PerplexityBot): the index assistants draw on when they answer with links. Googlebot's equivalent for AI answers.
  • User-initiated visits (ChatGPT-User, Claude-User, Perplexity-User): a user asks the assistant to read a page, or it checks a fact live. OpenAI and Perplexity state these agents do not crawl the web automatically.

Reference Table: The Main AI Crawlers

A summary of each vendor's official documentation, checked in September 2026. The robots.txt column uses their own wording.

Crawler (robots.txt token)VendorJobFollows robots.txt?If you block it
GPTBotOpenAITrainingYesContent signalled as off-limits for training; ChatGPT search unaffected
OAI-SearchBotOpenAIChatGPT searchYesOut of ChatGPT search answers (at most a navigational link)
ChatGPT-UserOpenAIUser actions in ChatGPT and Custom GPTsRules "may not apply"Not guaranteed
ClaudeBotAnthropicTrainingYesFuture content excluded from training
Claude-SearchBotAnthropicClaude searchYesReduced visibility in its results
Claude-UserAnthropicFetching a page for a userYesClaude cannot read your pages for its users
PerplexityBotPerplexitySearch (not used for foundation model training)YesNo longer surfaced or linked in Perplexity
Perplexity-UserPerplexityFetching a page for a user"Generally ignores" the rulesNot guaranteed
GooglebotGoogleGoogle Search, AI Overviews and AI Mode includedYesOut of Google, AI features included
Google-ExtendedGoogleControl token: Gemini training and grounding in the Gemini app and Vertex AIYesNo effect on Google Search; excluded from Gemini app grounding
Applebot-ExtendedAppleControl token: training Apple's modelsYesPages stay in Spotlight, Siri and Safari
CCBotCommon CrawlOpen web archive, heavily used for trainingYesFuture pages left out of the archive
meta-externalagentMetaModel training or indexing for Meta productsYes (a disallow blocks it)Excluded from both uses
bingbotMicrosoftBing index, reused in CopilotYesOut of Bing; NOARCHIVE is enough to leave Copilot answers only

Also worth knowing: Meta documents Meta-WebIndexer (Meta AI search) and Meta-ExternalFetcher (links requested by a user, which "may bypass" robots.txt). Applebot follows your Googlebot rules when it is not named. And according to Common Crawl, its archive supplies an estimated 70 to 90% of the training tokens behind nearly all large language models.

Blocking Training Is Not Leaving AI Answers

OpenAI puts it plainly: each setting is independent, and a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot for training. Anthropic separates ClaudeBot, Claude-SearchBot and Claude-User in the same way.

Google draws the line elsewhere. According to Google, Google-Extended has no effect on inclusion in Google Search and is not a ranking signal; it covers Gemini training, but also grounding answers in the Gemini app. For AI Overviews and AI Mode, Google says AI is integral to Search, so Googlebot is the control. Blocking Google-Extended therefore does not take you out of AI Overviews. To get out, Google has offered, worldwide since 31 August 2026, the Search generative AI control in Search Console (Settings): it excludes your site from AI Overviews, AI Mode and generative AI features in Discover without affecting rankings in regular Search. The nosnippet, max-snippet and noindex directives remain broader controls that also affect your regular results.

Apple's Applebot-Extended refuses training without touching Spotlight, Siri or Safari. At Microsoft, the NOARCHIVE meta tag keeps a page out of Bing Chat answers (Bing Chat is now Copilot) and out of training, while it stays in Bing.

Key takeaway: refusing training is a legitimate editorial choice. Refusing search crawlers and user visits is a visibility choice, and it is almost always accidental.

The robots.txt Grouping Rule That Trips Everyone Up

A robots.txt file is a series of groups: one or more User-agent lines, then Allow and Disallow rules. The standard (RFC 9309), Google and Bing agree on the essentials:

  • A crawler obeys a single group, the one that names it most specifically. A named group therefore replaces the * group rather than adding to it.
  • Crawler names ignore case; paths do not: /Admin/ and /admin/ are different.
  • The most specific rule wins; on a tie, Allow wins.
User-agent: *
Disallow: /admin/
Disallow: /cart/

# Mistake: this group replaces the * group for OAI-SearchBot,
# which may now crawl /admin/ and /cart/.
User-agent: OAI-SearchBot
Allow: /

And as the standard itself says, these rules are not a form of access authorisation. Confidential pages need a password.

Three Annotated robots.txt Setups

Use your own sitemap URL and allow about 24 hours for changes to apply: that is the delay OpenAI and Perplexity give, and Google generally caches robots.txt for up to 24 hours.

Setup 1: allow everything

# Every crawler, AI included, may read the public site.
User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://www.your-site.com/sitemap.xml

Our default for a company or services website: no crawler is named, so they all follow the same group.

Setup 2: allow search, refuse training

# Shared rules: search and user-initiated visits allowed.
User-agent: *
Allow: /
Disallow: /admin/

# Model training: refused.
# These crawlers get their own group and no longer follow the * group.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /

# Decide before uncommenting:
# - Google-Extended also covers grounding answers in the Gemini app;
# - meta-externalagent is also used to index content for Meta products.
# User-agent: Google-Extended
# User-agent: meta-externalagent
# Disallow: /

Sitemap: https://www.your-site.com/sitemap.xml

Search crawlers are not named, so they follow the * group and stay allowed.

Setup 3: block AI crawlers (not advised for an SMB)

User-agent: *
Allow: /
Disallow: /admin/

# Main documented AI crawlers (review this list quarterly): training, search and user-initiated.
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Meta-WebIndexer
User-agent: Meta-ExternalFetcher
Disallow: /

Two limits: ChatGPT-User, Perplexity-User and Meta-ExternalFetcher do not commit to following these rules, and Googlebot is untouched, so your pages remain eligible for AI Overviews unless you opt out in Search Console's Search generative AI control. Best kept for publishers whose content is the product (news, paid research).

The Traps That Block AI Without You Knowing

Blocking /_next/, scripts or stylesheets

Disallowing /_next/ (Next.js), /wp-includes/ (WordPress) or .js and .css files stops Googlebot rendering your pages properly. Google asks you not to block these resources when they help it understand the page, and Googlebot is what feeds AI Overviews.

A firewall or CDN that filters AI crawlers

A network-level block comes before robots.txt: the crawler gets a 403 and never reads your rules. Since July 2025, Cloudflare has asked every new domain at sign-up whether to allow AI crawlers. Since 15 September 2026, it offers three independent settings (search, training, agents) and, when a new domain is added, recommended settings based on the business model: for an ad-supported site, AI training is refused through robots.txt (Disallow AI Training) and AI agents are blocked on pages that display ads. Those agents are the visits your future clients trigger from ChatGPT or Claude. Beware too: choosing Block (or Block on pages with ads) for training now also blocks Googlebot, Bingbot and Applebot, and therefore search; to refuse training while keeping search, Cloudflare recommends Disallow AI Training. Check these settings after every migration.

Email addresses hidden behind a script

Cloudflare's Email Address Obfuscation, on by default, replaces each address with a link decoded in JavaScript. Researching this article with a tool that, like most AI crawlers, does not run JavaScript, we found one official documentation page whose contact address read "[email protected]". At least exempt your contact and legal pages.

Content that only exists in JavaScript

A Vercel analysis (December 2024) found that none of the major AI crawlers observed (OpenAI, Anthropic, Meta, ByteDance, Perplexity) rendered JavaScript; only Gemini, through Googlebot, and Applebot did. If your prices or FAQ only appear after a script runs, those crawlers see an empty page. Check the HTML your server sends:

# Is the key text in the HTML the server sends, without JavaScript?
curl -s https://www.your-site.com/pricing | grep -i "your key sentence"

Modern Next.js sites render pages on the server by default, and so does WordPress, which generates its pages in PHP — as long as the theme doesn't load key content through scripts.

Subdomains

Each host needs its own robots.txt: blog.your-site.com does not inherit from your-site.com, as Anthropic and Google both point out.

How to Check Your Server Logs

Only your access logs (host or CDN) show what crawlers actually do.

  1. Count visits per crawler (command below).
  2. Read status codes: 200 means served; 403 or 429 means a firewall or rate limit is blocking them.
  3. Verify identity: user agents can be faked. OpenAI, Anthropic, Perplexity and Common Crawl publish their crawlers' IP addresses, and Google documents a reverse DNS check.
  4. Track citations: in Search Console, the Search generative AI performance reports, available worldwide since 31 August 2026, measure your impressions in AI Overviews and AI Mode; the AI Performance report in Bing Webmaster Tools, in preview since February 2026, counts citations of your pages in Copilot, Bing's AI summaries and selected partner services.
# Visits per AI crawler in an access log
grep -ioE "gptbot|oai-searchbot|chatgpt-user|claudebot|claude-searchbot|claude-user|perplexitybot|perplexity-user|ccbot|meta-externalagent|bingbot" access.log \
  | tr 'A-Z' 'a-z' | sort | uniq -c | sort -rn

Google-Extended and Applebot-Extended will never appear: they are control tokens, not visitors.

Our Recommendation for an SMB That Wants AI to Recommend It

Let through everything that can get you cited with a link, and treat training separately:

  • Allow search crawlers: Googlebot, bingbot, Applebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot.
  • Allow user-initiated visits: ChatGPT-User, Claude-User, Perplexity-User. These are your prospects asking an assistant to read your page.
  • Decide on training deliberately. No vendor documentation links blocking GPTBot or ClaudeBot to lower visibility in answers, or the reverse. For a company website we usually leave it open; for content with its own market value, refusing is legitimate. Keep Google-Extended open if the Gemini app matters to you.

Access is only the first step: next come citable content, consistent information about your business and, as a complement, an llms.txt file. That file is still a proposal by Jeremy Howard (2024), not an official standard, and Google says no special file is needed for its AI features: see our article on the llms.txt file.

Checklist: your AI crawlers in 10 points

  • No Disallow: / aimed at Googlebot, bingbot, OAI-SearchBot, Claude-SearchBot or PerplexityBot.
  • Every named group repeats the relevant rules from the * group.
  • No block on /_next/, .js or .css files.
  • An explicit choice for GPTBot, ClaudeBot, CCBot, Applebot-Extended, Google-Extended and meta-externalagent.
  • CDN or firewall AI settings reviewed.
  • Contact and legal pages exempt from email obfuscation.
  • Prices, services, address and FAQ present in the served HTML.
  • One robots.txt per subdomain.
  • Logs reviewed 24 to 48 hours after each change.
  • A quarterly review: new crawlers, updated documentation.

What Agence Zen does. We treat crawler access as a technical prerequisite of GEO: a robots.txt written by use case, server-rendered pages (Next.js or WordPress with a hand-coded theme), CDN and obfuscation settings checked, then structured data and llms.txt. GEO foundations (structured data, llms.txt, FAQ written to be cited) are included from the Essentiel package, full GEO in Signature, and our SEO & GEO retainers add monthly tracking of what AI assistants say about your business. Everything happens remotely, by video call. Explore our SEO & GEO services.

FAQ

Should I block GPTBot?

Only if you do not want your content used to train OpenAI's models. According to OpenAI, GPTBot is used for training, while ChatGPT search relies on OAI-SearchBot, which is controlled independently. Blocking GPTBot does not stop you appearing in ChatGPT search, provided OAI-SearchBot is allowed. For an SMB seeking visibility, it is not a requirement.

Does blocking Google-Extended remove me from AI Overviews?

No. Google states that Google-Extended does not affect inclusion in Google Search and is not a ranking signal. It covers training of future Gemini models and grounding in the Gemini app and Vertex AI. To leave AI Overviews and AI Mode without leaving Google, use the Search generative AI control in Search Console, available worldwide since 31 August 2026; it does not affect rankings in regular Search.

Do all AI crawlers follow robots.txt?

No. According to their documentation, GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, Claude-User and PerplexityBot follow it. OpenAI says the rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because a user triggered the visit. robots.txt is an instruction, not a lock.

Why is my site missing from ChatGPT when my robots.txt allows everything?

The block often sits elsewhere: a firewall or CDN filtering AI crawlers, content that only appears through JavaScript, an email address hidden behind a script, or simply a lack of authority. Check your server logs first. If OAI-SearchBot receives 403 responses, the problem is technical. If it reads your pages, the work lies in content and reputation.

What setup do you recommend for an SMB?

Allow search crawlers (Googlebot, bingbot, Applebot, OAI-SearchBot, Claude-SearchBot, PerplexityBot) and user-initiated agents (ChatGPT-User, Claude-User, Perplexity-User): they are what gets you cited with a link. Training is a business decision. For a company website we usually leave it open; for paid or exclusive content, refusing it is legitimate.

Sources

Can AI assistants actually read your site?

Let's talk about your visibility in AI answers on a 30-minute discovery call, by video, with the founder.

Book a discovery call→
30-min discovery call
Reply within 1 business day
Detailed quote, fixed price