FirmAEO
Engines · Per-engine retrieval

ChatGPT, Gemini, Perplexity, Claude and Copilot: which index each one reads

David TerrellFounder, Firm AEOSeptember 23, 202611 min read

Five engines, three kinds of plumbing. Google’s AI Overviews and AI Mode retrieve from the Google Search index, which Google states in its own documentation. Microsoft Copilot retrieves from the Bing index, which Microsoft states in its own documentation. ChatGPT, Perplexity and Claude each run documented retrieval crawlers of their own, and none of the three publishes what its index is made of.

That split is the whole planning problem. Firm AEO has already published which directories each engine cites and why the engines disagree about which firm to name. This post takes the layer underneath both: the substrate each engine retrieves from, which of the popular claims about that substrate a vendor has actually confirmed, which crawler shows up in a law firm’s server logs, and what a firm can do with the answer. The short version is that per-engine optimization is mostly a myth, per-engine measurement is mostly unavailable, and the one engine with a first-party dashboard is the one with the least reach among people looking for a lawyer.

Which AI engines should a law firm optimize for?

ChatGPT first, Google’s AI surfaces alongside it, and the other three as coverage rather than as separate programs. The ordering comes from reach. In iLawyerMarketing’s August 17, 2026 survey of 1,110 US adults, 41.9% said they would use ChatGPT to research which lawyer to hire, against 15.1% for Gemini, 6.2% for Claude and 3.7% for Perplexity. Google itself still led every channel at 71.9%.

Reach is not the same as willingness to name a firm, and the two pull in different directions. SOCi’s 2026 Local Visibility Index, published February 17, 2026 across 2,751 brands and roughly 350,000 locations, found ChatGPT recommending a specific location in 1.2% of cases, Perplexity in 7.4% and Gemini in 11%, against 35.9% for Google’s local three-pack. The engine most people say they would ask is the engine least likely to hand back a business name.

The third column below is the one that changes the work. Martindale-Avvo’s April 6, 2026 analysis reported that ChatGPT’s legal answers match Google’s top ten results less than 25% of the time, while Perplexity and Claude match at roughly 75% and Gemini at roughly 50%. Ranking work carries into three of these engines fairly well and into ChatGPT poorly.

Three separate samples, placed side by side for planning only. Reach is iLawyerMarketing (n=1,110 US adults, August 2026), naming rate is SOCi (February 2026, all local business categories rather than legal), overlap is Martindale-Avvo (April 2026, legal queries). The columns are not drawn from one study and should not be read as a single ranking.
EngineWould use it to research a lawyerNames a specific local businessOverlap with Google’s top ten
ChatGPT41.9%1.2% of locationsUnder 25%
Gemini15.1%11% of locationsAbout 50%
Claude6.2%Not measuredAbout 75%
Perplexity3.7%7.4% of locationsAbout 75%
Microsoft CopilotNot measured separatelyNot measuredNot measured
Google Search, for comparison71.9%35.9% in the local three-packNot applicable

What does each engine actually retrieve from?

Two of the five engines have a vendor statement about their retrieval substrate, and three do not. Google Search Central documents that a page must be indexed and eligible to be shown with a snippet before it can appear as a supporting link in an AI Overview or AI Mode. Microsoft documents that its Copilot experiences draw on content Bingbot has crawled. OpenAI, Perplexity and Anthropic each publish crawler documentation without publishing index composition, which means every confident sentence about what ChatGPT runs on is a third-party inference.

The crawler column is the part a law firm can verify without asking anyone. Each of these user agents is named in the vendor’s own documentation, and each one shows up in raw server logs. A firm whose logs contain Googlebot and Bingbot but never OAI-SearchBot, PerplexityBot or Claude-SearchBot has a retrieval problem it can see for itself, months before it shows up as a missing answer. Whether to allow or restrict any of these agents is a separate decision, and Firm AEO treats the training crawlers and the search crawlers as different questions.

Retrieval substrate by engine, with the vendor documentation status for each claim. Crawler names are taken from each vendor’s published crawler documentation, September 2026.
EngineRetrieval substrateVendor confirmed?Crawler in your logsFirst-party reporting
Google AI Overviews and AI ModeThe Google Search indexYes. Google states index and snippet eligibility are the preconditionGooglebot, with Google-Extended governing Gemini model useSearch Console, with no AI citation breakout
Microsoft CopilotThe Bing indexYes. Microsoft documents Copilot experiences using Bingbot-crawled contentBingbotBing Webmaster Tools AI Performance report
ChatGPTOpenAI retrieval plus licensed and partner contentPartly. OpenAI documents a dedicated search crawler and does not publish index compositionOAI-SearchBot, ChatGPT-User, GPTBotNone
PerplexityPerplexity’s own crawl plus partner contentPartly. Perplexity documents its crawlers and does not publish index compositionPerplexityBot, Perplexity-UserNone
ClaudeAnthropic search retrieval plus third-party search resultsPartly. Anthropic documents three crawlers and does not publish index compositionClaude-SearchBot, Claude-User, ClaudeBotNone

Why does ChatGPT disagree with Google more than the other engines do?

Because ChatGPT leans hardest on sources that sit outside the ranked result set. Martindale-Avvo’s April 2026 figure of under 25% overlap with Google’s top ten is the widest gap of the five engines, and the source mix explains the direction of it. Profound data reported through Claude for Lawyers puts Wikipedia at roughly 47.9% of ChatGPT’s top-ten source share, and Reddit at roughly 46.7% of Perplexity’s. Neither of those is a page a law firm ranks for.

The practical reading is diagnostic. A firm that holds position one on Google and still never appears in ChatGPT does not have a ranking problem. It has a corroboration problem, which means the firm is not described consistently enough, in enough independent places, for a retrieval system to assemble it into an answer. Firm AEO’s post on how AI decides which law firm to cite carries the corroboration model in full, and this post assumes it rather than repeating it.

The reverse case is just as common and reads as an easier fix. A firm that appears in Perplexity and Claude but not in Google’s AI Overviews is usually running into eligibility rather than reputation, because those two engines overlap Google’s top ten at roughly 75% while Google requires index and snippet eligibility before an AI Overview can cite anything at all. Ahrefs found in March 2026 that only about 38% of AI Overview citations rank in Google’s top ten, so the eligibility gate is not a ranking gate, but it is still a gate.

Does work done for one engine carry over to the others?

Most of it does, which is why per-engine programs are usually a way to be sold the same work five times. The engines differ in what they retrieve from and agree almost completely on what makes a source usable: a page that can be crawled, a business that resolves to one identity, a sentence that survives being quoted alone, and a claim that something other than the firm’s own website also says. The Princeton generative engine optimization study by Aggarwal and colleagues, presented at KDD 2024, measured visibility lifts of roughly 22% to 41% from adding statistics, citations and quotations, and measured them across engines rather than inside one.

Four assets do the carrying, and a law firm can audit all four in an afternoon:

  • One resolvable entity. Firm name, address, phone and practice descriptions identical across the website, Google Business Profile, Bing Places and every directory profile, with schema that matches the visible text rather than contradicting it.
  • Retrievable pages. Indexed, snippet-eligible, and not blocked to the search crawlers each vendor documents by name.
  • Extractable sentences. A 40 to 75 word answer at the top of every practice page that stays accurate with no surrounding context, because that is the unit an engine lifts.
  • Independent corroboration. Directory profiles, bar records, press and review platforms that describe the firm the same way, which is the only input a firm cannot manufacture on its own site.

Which engine is most likely to name a law firm at all?

The differences between engines are smaller than the differences between questions. Citorian’s five-engine study of personal injury prompts, run June 13 to 15, 2026 across 359 answers in five US metros, found that the engines named at least one firm 62% of the time overall, and that all five sat within about ten points of each other, from Gemini at the top of the range to Google’s AI Overviews at the bottom. Firm AEO’s personal injury post carries the per-engine table and the most-named firms.

Phrasing moved the result far more than engine choice did. In the same dataset, direct vetting prompts produced a firm name most of the time while situational prompts, which is how a frightened person actually types, produced one in under a fifth of answers. A firm choosing between engines is optimizing a ten-point spread. A firm choosing which questions to answer on its own site is working on a spread several times that size, and Firm AEO treats that as the larger lever.

One more caution belongs here. The SOCi and Citorian numbers measure different things, and neither is a forecast for a particular firm. SOCi measured how often any specific location gets recommended across all local business categories. Citorian measured how often a personal injury prompt in a large metro produced any firm name. Neither number tells a family law practice in a mid-sized city what to expect, and any agency presenting either one as a projected outcome is misreading its own source.

Where do licensing deals put a law firm out of crawl range?

Some answer surfaces are entered by qualifying, not by publishing, and no amount of on-site work reaches them. Best Lawyers launched a ChatGPT app on April 22, 2026 that answers requests of the form find me a specialty lawyer in a location using its peer-review rankings. LegalZoom announced a partnership with Perplexity on June 4, 2025 that places LegalZoom services and subscriber discounts inside Perplexity answers. Neither surface is reachable by improving a website.

Firm AEO treats these as a separate line item from retrieval work, with two rules. The first is that inclusion in a peer-review directory follows the directory’s own process, on the directory’s own timeline, and a marketing vendor that promises placement is describing something it does not control. The second is a bar-rule limit: ABA Model Rule 7.1 forbids false or misleading communications about a lawyer’s services, so a listing earned through a ranking process may be stated accurately and may not be restated as a claim of superiority the process does not support. The firm’s ethics counsel decides what may be published, not the marketing vendor.

Which engine can a law firm actually measure?

Exactly one of the five reports its own citations, and it is not the one with the reach. Microsoft shipped an AI Performance report inside Bing Webmaster Tools in February 2026 showing grounding queries and citations across Copilot and Bing surfaces. It remains the only first-party citation dashboard any of these vendors offers, and almost nobody in legal marketing is reading it, which makes it both the cheapest measurement available and the least crowded.

Everything else is inference, and the honest measurement plan says so. Firm AEO runs it in this order:

  • Bing Webmaster Tools AI Performance, monthly. First-party citation counts for Copilot and Bing, and the only number in this list that is not an estimate.
  • Server logs, monthly. Presence and frequency of OAI-SearchBot, PerplexityBot, Claude-SearchBot and Googlebot, read as evidence of retrieval rather than of citation.
  • Google Search Console, monthly. Index coverage and snippet eligibility, which Google states is the precondition for an AI Overview citation, and which Search Console reports even though it does not break out AI citations.
  • A logged-out prompt panel, monthly. A fixed set of prompts per metro and practice area, run three times each in a signed-out session, scored for whether the firm is named. Weekly runs mostly measure model variance.
  • GA4 referrals, monthly and with a known ceiling. A large share of assistant traffic arrives with no referrer, so this number reads as a floor rather than a total.

Where should a law firm spend first?

Spend on the substrate every engine shares before spending on any engine by name. The ordering below is Firm AEO’s judgement about sequence, not a measured ranking, and the reason for it is arithmetic rather than preference: four of the five engines cannot be measured directly, so work whose value depends on knowing which engine responded is work a firm cannot verify it bought.

A typical first 90 days runs in this order:

  • Weeks 1 to 2. Entity reconciliation across the website, Google Business Profile, Bing Places and the major legal directories, plus a crawler audit that confirms the documented search agents are not blocked.
  • Weeks 2 to 6. Directory profile completion and correction, which supplies most of the legal citation layer in every published dataset.
  • Weeks 3 to 10. Answer-first rewrites of the practice pages, one extractable answer per question a client actually asks, with a source for every number.
  • Weeks 4 onward. A review program that respects Model Rule 7.1 and the Federal Trade Commission rule on fake reviews, which took effect October 21, 2024.
  • Month 2 onward. Baseline measurement across the five surfaces above, with the first meaningful comparison at month three rather than month one.
Frequently asked

Which AI engine sends the most traffic to law firm websites?

ChatGPT, by a wide margin, though the total remains small. Previsible’s July 6, 2026 report covering 6.77 million AI sessions across 166 properties put ChatGPT at 92.4% of standalone large language model referrals. The total those referrals add up to is still modest in legal: Martindale-Avvo put AI tools under 5% of law firm website traffic in April 2026, and Great Jakes measured AI referrals at 0.47% of sessions across its law firm client portfolio.

Does Microsoft Copilot use the same index as ChatGPT?

Microsoft documents Copilot as grounded in the Bing index. OpenAI does not publish what ChatGPT’s search retrieval is composed of, so the widely repeated claim that ChatGPT runs on Bing is a third-party description rather than a current vendor statement. The practical answer does not depend on settling it: Bing Places and Bing Webmaster Tools are inexpensive, and one of the two engines definitely reads that index.

Should a law firm optimize differently for Claude?

Not materially. Claude’s legal answers overlapped Google’s top ten at roughly 75% in Martindale-Avvo’s April 2026 analysis, which is the closest of the five engines to Perplexity and the furthest from ChatGPT. A firm doing sound indexation, entity and corroboration work is already doing most of what Claude rewards. The one Claude-specific step is confirming that Claude-SearchBot is not blocked at the server or the content delivery network.

Is Google AI Mode a different engine from AI Overviews?

They are two surfaces of the same index. Google Search Central applies the same eligibility rule to both: a page must be indexed and eligible for a snippet before it can appear as a supporting link. AI Mode decomposes a question into more sub-queries than an AI Overview does, which changes how much narrow sub-practice coverage pays, and Firm AEO treats that as a content-depth question rather than a separate optimization track.

Do we need separate pages or a separate site for each AI engine?

No. Google Search Central states that appearing in its AI features requires no additional machine-readable files, no AI-specific structured data and no special optimizations, and no other vendor has published a requirement of that kind either. Duplicating pages per engine creates the near-duplicate content that Google’s scaled content abuse policy targets, which makes it a risk rather than a hedge.

Find out which engines can see your firm, and which cannot.

Firm AEO runs the five-engine baseline: server-log retrieval evidence, Bing AI Performance citations, and a logged-out prompt panel for your metro and practice areas. You get the scorecard and the order of work, not a per-engine invoice.