How AI Crawlers Work
AI crawlers are the quiet infrastructure behind answer engines, training datasets, and AI search results. Here is how they discover, fetch, and reuse web content — and how publishers control what they take.
Every AI answer that cites the web started as a crawl. Before a chatbot can quote your article, a bot had to find it, download it, and file it away. These bots — AI crawlers — now account for a large share of web traffic, and they do three different jobs that are often confused: building training datasets, maintaining search indexes, and fetching pages live to answer a user's question. Understanding the difference is the key to controlling them.
The three jobs AI crawlers do
- Training crawls collect text to build or improve AI models. Example user agents: GPTBot (OpenAI), ClaudeBot (Anthropic), Meta-ExternalAgent. Blocking these opts your content out of future model training — it does not affect live answers.
- Search-index crawls build the indexes that AI search products retrieve from. Examples: OAI-SearchBot, Claude-SearchBot, PerplexityBot. Blocking these removes your pages from AI-generated answers.
- User-triggered fetches happen when someone pastes your link into a chatbot and asks about it. Examples: ChatGPT-User, Claude-User, Perplexity-User. Blocking these stops assistants from opening and quoting your pages on demand.
This three-way split is the single most important concept in crawler control. A publisher who wants to be cited in AI answers but not used for training should allow the search bots and block the training bots — and many sites accidentally block everything, or nothing, because they never made the distinction.
How a crawl works, step by step
1. Discovery. The crawler starts from known URLs — sitemaps you publish, links from already-crawled pages, and feeds — and builds a frontier of pages to visit. A clear XML sitemap and a sensible internal link structure are still the fastest way to get discovered, by AI crawlers no less than by Googlebot.
2. Permission check. Well-behaved crawlers fetch your robots.txt first and obey its directives. This file is the entire legal and technical basis of crawler control: it is voluntary, but the major AI companies honor it, and ignoring it would destroy their relationships with publishers.
3. Fetch. The bot requests the page over HTTP, identifies itself with a user-agent string, and downloads the HTML. Most AI crawlers execute little or no JavaScript — content that only renders client-side is often partially or entirely missed. Server-rendered HTML remains the most reliably ingestible format on the web.
4. Extract and store. The crawler strips navigation, ads, and boilerplate, keeps the main text, and records metadata: URL, title, publication date if it can find one, and language. Duplicates and near-duplicates are collapsed so the same article syndicated across sites does not count many times.
5. Re-crawl. Popular and frequently updated pages are revisited on a schedule; obscure pages may wait months. How often a crawler returns depends on the page's apparent importance and change frequency — another reason accurate last-modified signals help.
The main AI crawlers and their user agents
- OpenAI: GPTBot (training), OAI-SearchBot (search index), ChatGPT-User (fetches pages when a user asks). OpenAI publishes its IP ranges so publishers can verify genuine crawlers.
- Anthropic: ClaudeBot (training), Claude-SearchBot (search), Claude-User (user-triggered fetch). Anthropic asks publishers to add rules on every subdomain, supports crawl-delay directives, and warns that IP-blocking its bots can break the opt-out, since the bot must be able to read robots.txt.
- Perplexity: PerplexityBot (indexing) and Perplexity-User (on-demand fetch). Perplexity answers cite sources heavily, so allowing these bots is the direct path to being cited in its answers.
- Google: Google-Extended (the opt-out specifically for AI training and model features — Google states it does not affect Search ranking), alongside regular Googlebot, which increasingly feeds AI Overviews.
- Meta: Meta-ExternalAgent (training and AI features) and Meta-ExternalFetcher (user-triggered fetches, e.g. link previews).
- Apple: Applebot (used across Siri, Spotlight, and Apple Intelligence features) and Applebot-Extended (training opt-out).
- Others: Bytespider (ByteDance), CCBot (Common Crawl — an open dataset used by many models and researchers; blocking it is the broadest single opt-out), and DuckAssistBot (DuckDuckGo's answer features).
User-agent strings change, so treat any list — including this one — as a snapshot. Check each vendor's published crawler documentation before writing rules, and re-check quarterly.
What crawlers can and cannot read
AI crawlers are excellent at plain HTML and mediocre at everything else. Text in the main document is captured; text injected by JavaScript after load often is not. Images are typically noted, not understood, unless the crawler runs a vision model over them. Video and audio transcripts depend on whether captions or transcripts are published in the HTML. Paywalled content is fetched only if the crawler is granted access — which is why licensing deals, not scraping, govern most premium content in AI answers.
The practical lesson for publishers: if you want AI systems to read and cite a page, serve its core content as server-rendered HTML with a clear title, publication date, and author. Everything else — schema markup, llms.txt files, clean Markdown alternates — is optimization on top of that foundation.
How publishers control crawlers with robots.txt
Control happens in the robots.txt file at your site's root. The most common deliberate setup is: visible in AI answers, out of training data. It looks like this:
Allow answer and search crawlers, block training crawlers:
- Allow: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot — these power citations and live answers.
- Disallow: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, Bytespider, CCBot — these feed training.
The most common accident is a copied "block all AI" snippet that includes the search bots too — for example, disallowing GPTBot, OAI-SearchBot, and ChatGPT-User together. That was probably meant as a training opt-out, but it also removes the site from ChatGPT search and stops ChatGPT opening its links. Split the groups by job, not by company.
Two caveats. First, robots.txt is advisory — compliance is voluntary, though the major labs honor it. Second, your CDN, firewall, or bot-protection layer can silently discard crawler requests even when robots.txt allows them; studies of large sites have found household names accidentally invisible to AI crawlers because of exactly this. Check server logs for the user agents above to confirm crawlers actually reach your pages.
Server load and crawler politeness
AI crawlers can be aggressive. Training crawls in particular fetch at high concurrency, and smaller sites have reported traffic spikes that look like denial-of-service events. Defenses include crawl-delay directives (honored by some bots, ignored by others), rate limiting at the server or CDN, and IP-range allowlists that admit only verified crawler IPs — which also blocks impostor bots spoofing user-agent strings. If crawler traffic is hurting performance, the fix is traffic shaping first and blocking second; a blanket block also forfeits any visibility the crawler provided.
What this means for readers
When an AI assistant says it cannot access a page, the cause is often one of these crawler mechanics: the site blocks AI user agents, the content requires JavaScript the fetcher does not run, or the page sits behind a login or paywall. It is not always a knowledge cutoff. Knowing this helps you judge whether "I can't browse that" is a real limitation or a site configuration — and whether pasting the article's text directly will get you the analysis you wanted.
Checklist: audit your site's AI crawler policy
- Read your robots.txt and list every AI user agent it mentions. If it mentions none, everything is allowed by default.
- Decide per category: training (usually block or license), search index (usually allow), user fetch (usually allow).
- Check server logs for GPTBot, ClaudeBot, PerplexityBot, and OAI-SearchBot hits to confirm your rules work in practice — and that your firewall is not silently blocking allowed bots.
- Verify crawler IPs against vendors' published ranges to distinguish genuine crawlers from spoofed scrapers.
- Serve citable HTML: server-rendered content, clear titles, dates, and authorship on every page you want cited.
- Review quarterly. User agents, vendor policies, and your own business stance all change; a robots.txt file from 2024 is probably wrong in 2026.
FAQ
Does blocking GPTBot remove my site from ChatGPT answers?
No. GPTBot is OpenAI's training crawler. Blocking it opts your content out of model training. To leave or enter ChatGPT's live answers, the relevant bots are OAI-SearchBot (indexing) and ChatGPT-User (on-demand fetch).
Does Google-Extended affect my Google Search rankings?
According to Google, no. Google-Extended is specifically the opt-out for AI training and related model features, and Google states it has no effect on Search ranking or inclusion.
Can AI crawlers read JavaScript-heavy pages?
Mostly no. The majority of AI crawlers fetch raw HTML and execute little or no JavaScript. Content that only appears after client-side rendering is frequently missed. Server-rendered HTML is the safest format.
Is robots.txt legally binding?
It is a voluntary convention, not a law — but the major AI companies honor it, and courts and regulators increasingly treat deliberate circumvention as evidence of bad faith. For practical purposes, compliant crawlers obey it.
Is this official Grokipedia documentation?
No. GrokExpedia is an independent educational publication and is not affiliated with xAI, Grok, Grokipedia, Wikipedia, or Wikimedia Foundation.