A law firm should allow the AI search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Perplexity-User, Claude-User and Googlebot) if it wants its pages cited in ChatGPT, Claude, Perplexity and Google AI answers, and can decide separately whether to block the training-only bots (GPTBot, ClaudeBot and Google-Extended) without affecting that citation eligibility. The two questions get treated as one in most generic robots.txt guides, and they are not one question.
Each of the three companies that publish a training crawler also publishes at least one differently named bot for search or live retrieval, and each company documents different behavior for whether robots.txt actually controls the live-fetch bot. A firm that copies a one-size disallow block from a generic SEO checklist can end up blocking the exact bot that would have cited it.
- GPTBot, ClaudeBot and Google-Extended are training-only bots. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Perplexity-User, Claude-User and Googlebot are the bots that actually retrieve pages for citation in ChatGPT, Claude, Perplexity and Google AI answers.
- Blocking a training bot does not remove a firm from citation eligibility in that company's answer engine, because each company runs training and search crawling as separate systems under separate robots.txt user-agent names.
- Ahrefs' analysis of robots.txt files across roughly 140 million websites, published May 21, 2025, found GPTBot blocked by 5.89% of sites, ClaudeBot by 5.74% and PerplexityBot by 5.61%. Explicit, bot-specific disallow rules were far rarer, at 0.5% for GPTBot.
- Compliance is uneven across companies. Anthropic states that all three of its bots, including the live user-fetch bot, honor robots.txt. OpenAI and Perplexity both say their user-initiated fetch bots, ChatGPT-User and Perplexity-User, may ignore robots.txt because a person, not an automated crawler, triggered the request.
- AI training crawlers grew from 22% of AI crawler requests on Cloudflare's network in spring 2025 to 52% in June 2026, according to Cloudflare's September 30, 2026 report, so the crawler traffic a firm sees in its server logs is shifting toward training bots even though citation depends on the search bots.
What is the difference between a training bot and a search bot?
A training bot collects pages to improve a future model. A search or retrieval bot fetches pages to answer a live question, whether that question comes from an automated index build or from a specific person typing a prompt. The distinction matters because blocking one has no effect on the other; they are different crawlers under different user-agent names, run by the same company for different purposes.
OpenAI's own documentation states this plainly: GPTBot is used for training, and "disallowing GPTBot indicates a site's content should not be used in training." A separate crawler, OAI-SearchBot, powers ChatGPT's search feature, and OpenAI says sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers." A third bot, ChatGPT-User, fetches a page only when an actual person asks ChatGPT a question that needs that page.
Anthropic and Perplexity split their crawlers the same way, under their own names, and Google splits the same function differently: Google-Extended is not a crawler at all. Google's own crawling documentation states that "Google-Extended is used by Gemini apps and Vertex AI API for Gemini," a training-and-grounding control token. Googlebot does the actual fetching for Search, and AI Overviews and AI Mode are Search features, not a separate product Google-Extended governs.
Which bots does a law firm need to allow to be cited?
The search and live-retrieval bots, not the training bots. A firm that wants to appear in a ChatGPT, Claude or Perplexity answer needs OAI-SearchBot, Claude-SearchBot and PerplexityBot to be able to crawl its site, and needs ChatGPT-User, Claude-User and Perplexity-User to be able to fetch a specific page when a user's question points to it. Googlebot has to be allowed for anything to show up in Google Search, AI Overviews or AI Mode, which is table stakes for any site already doing SEO.
The table below lists each company's bots by function, drawn from each company's own published documentation as of September 2026.
| Company | Training-only bot | Search / indexing bot | Live user-fetch bot | Does the live-fetch bot honor robots.txt? |
|---|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User | No — OpenAI says robots.txt rules may not apply, since a user triggered the fetch |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User | Yes — Anthropic says all three bots honor robots.txt, including Crawl-delay |
| Perplexity | Not separately named (PerplexityBot is search-only) | PerplexityBot | Perplexity-User | No — Perplexity says this fetcher "generally ignores robots.txt rules" |
| Google-Extended (training/grounding token, not a crawler) | Googlebot | Not applicable — Googlebot covers Search, AI Overviews and AI Mode | Governed by standard Googlebot rules |
Does blocking GPTBot, ClaudeBot or Google-Extended hurt a firm's AI visibility?
No, not for citation purposes, because citation runs through the search bots, not the training bots. A firm that disallows GPTBot in robots.txt is telling OpenAI not to use its pages to train a future model. It is not telling OAI-SearchBot to stay away, and OAI-SearchBot is the crawler that determines whether the firm's pages can be retrieved into a ChatGPT search answer today.
The same logic holds for Google-Extended. Disallowing it opts a site out of having its content used to train or ground Gemini apps and the Vertex AI API for Gemini. It does not touch Googlebot, and it does not remove a page from eligibility for AI Overviews or AI Mode, both of which are Search surfaces built on the regular Googlebot index.
Whether to block the training bots is therefore a separate, legitimate business question about how a firm's content is used to build a competitor's product, not a visibility question. A firm can block GPTBot and Google-Extended and still be fully eligible for citation in ChatGPT and Google AI answers, provided it leaves OAI-SearchBot, Googlebot and the other search bots untouched.
How many sites actually block AI crawlers, and does blocking work?
A small minority block by name, and even a declared block is not always honored the same way by every bot. Ahrefs analyzed robots.txt files across approximately 140 million websites in a study published May 21, 2025 and found GPTBot blocked (by any rule, including a blanket disallow that catches all bots) on 5.89% of sites, ClaudeBot on 5.74% and PerplexityBot on 5.61%. Deliberate, bot-specific disallow rules were far less common: 0.5% of sites for GPTBot, 0.28% for ClaudeBot and 0.15% for PerplexityBot.
Compliance is fragmented even where a block is declared. Perplexity's own documentation says Perplexity-User "generally ignores robots.txt rules" because a user, not an automated crawl, requested the fetch. OpenAI says the same of ChatGPT-User. Anthropic is the outlier: its documentation states that Claude-User, its equivalent live-fetch bot, still honors robots.txt along with ClaudeBot and Claude-SearchBot. A firm cannot assume a robots.txt rule controls a live-fetch bot unless it checks that specific company's documentation.
The volume behind all of this is growing fast. Cloudflare's September 30, 2026 report on its network traffic found that AI training crawlers made up 52% of AI crawler requests in June 2026, up from 22% in spring 2025, and that automated traffic overall, AI agents included, now exceeds half of everything Cloudflare handles, with AI agent requests up more than 1,700% year over year. Most of that growth is training and agent traffic a firm's robots.txt can meaningfully stop; the search bots that matter for citation are a smaller, separate share of it.
What should a law firm's robots.txt actually allow?
A rule that names each search and live-fetch bot explicitly, rather than a blanket line that happens to catch them by accident. The most common way a firm accidentally blocks its own citation eligibility is a wildcard Disallow rule added for an unrelated reason, such as blocking a staging subdomain, that also matches every user agent including the ones a firm needs.
- Confirm robots.txt has no blanket "Disallow: /" under a wildcard User-agent: * block that would catch OAI-SearchBot, Claude-SearchBot and PerplexityBot along with everything else.
- Add explicit Allow rules for OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User if the site's default rule is otherwise restrictive.
- Decide GPTBot, ClaudeBot and Google-Extended separately, as a training-data policy question for the firm's leadership, not as part of the citation checklist.
- Check that practice-area and attorney bio pages, the pages most likely to be cited, are not excluded by an older robots.txt rule written before AI bots existed.
- Re-check the file after any site migration or CMS change, since migrations are the most common point where an old blanket rule gets reintroduced.
If I block GPTBot, will ChatGPT stop recommending my firm?
No. GPTBot only controls whether a site's content can be used to train a future OpenAI model. Citation in a live ChatGPT answer depends on OAI-SearchBot and ChatGPT-User, which are separate bots under separate names in OpenAI's own documentation. A firm can block GPTBot and still be fully eligible for citation.
Do we also need an llms.txt file?
That is a separate question from robots.txt, and the two are often confused. No major AI engine has confirmed that it reads an llms.txt file for citation decisions; robots.txt is the mechanism every AI company documents and, per its own stated rules, actually checks.
Does blocking Google-Extended affect whether my firm shows up in AI Overviews?
No. Google's documentation describes Google-Extended as a training and grounding control for Gemini apps and the Vertex AI API for Gemini, separate from Googlebot. AI Overviews and AI Mode are Search features built on the regular Googlebot crawl, so Google-Extended has no bearing on whether a page can appear there.
How can a firm check what its current robots.txt actually allows?
Fetch the live file at the site's root, such as https://example.com/robots.txt, and read every User-agent block in it, not just the one labeled with a wildcard. A blanket rule written for an unrelated purpose, like keeping a staging environment out of Google, can catch AI search bots by accident if it is not scoped to the specific user agent it was meant for.
- 1.OpenAI, Crawlers and bots documentation
- 2.Anthropic (Claude support), Anthropic's web crawling bots
- 3.Perplexity, Perplexity crawlers documentation
- 4.Google, Crawling infrastructure overview
- 5.Ahrefs, AI bot block rates study (May 21, 2025)
- 6.Cloudflare, The internet has a second audience (September 30, 2026)
- 7.Seer Interactive, Do LLMs respect robots.txt?