FirmAEO
Measurement · Monthly protocol

How to track whether ChatGPT is mentioning your law firm

David TerrellFounder, Firm AEOSeptember 25, 20267 min read

Track it by running a fixed prompt set in logged-out sessions, three runs per prompt per engine, on the same week each month, and recording which firms every answer named and which sources it cited. The number that matters is the share of answers naming the firm across the whole set. One run is not a measurement.

Firm AEO has published the run-it-yourself steps before, and this post is not another list of them. It is about the instrument: how many answers the sample needs before a percentage means anything, which engines still earn a slot in the monthly run, what each row of the ledger has to record, and what a firm is allowed to say out loud about the result.

Key takeaways
  • Answer engines are not deterministic, so a single run is an anecdote. Citorian ran each prompt three times in logged-out sessions for its June 2026 five-engine study, which is the smallest repeat count that separates a real appearance from a lucky one.
  • Citorian logged 359 usable answers out of 375 query runs between June 13 and 15, 2026. Roughly one run in twenty-three returned an error or no answer, so a measurement protocol has to record failed runs rather than quietly drop them.
  • Sample size decides what the number can say. In a 15-answer sample one flipped answer moves the reported share by 6.7 points, which is larger than most real month-over-month movement.
  • The Bing Webmaster Tools AI Performance report, in public preview since February 9, 2026, is the only first-party citation dashboard any engine publishes. It reports total citations, average cited pages, page-level citation activity and grounding queries, and it carries no click data.

How do I track whether ChatGPT is mentioning my law firm?

By building a fixed prompt set and scoring appearance rate against it every month. The prompt set is written once and then frozen, because a set that changes between months measures the change in the prompts rather than the change in the firm’s visibility. Firm AEO treats the frozen set, the logged-out session and the fixed run count as the three conditions that make two months comparable.

Prompt mix decides the ceiling before anything else does. Citorian’s June 2026 five-engine test found that direct vetting asks produced a named firm 94% of the time, specific-injury asks 79%, question-led asks 23% and situational asks 18%. A prompt set made only of direct asks will report a flattering number that no ordinary client ever triggers, and a set made only of situational asks will report near zero for a firm that is doing well.

Why does the session have to be logged out, and why three runs?

Logged out to strip personalization, three runs to catch variance. Citorian used logged-out sessions specifically to limit personalization, and ran each query three times to capture variance. A signed-in account carries chat history, saved memory and account-level location, which means a signed-in test measures the relationship between the engine and the person running the test.

The failed runs matter as much as the successful ones. Citorian collected 359 logged answers from 375 query runs between June 13 and 15, 2026, with the gap explained as sessions that returned errors or no answer. That is roughly one run in twenty-three producing nothing. A ledger that silently drops those runs inflates every percentage computed from it, so Firm AEO records a failed run as its own outcome alongside answers that named no firm at all.

How many answers does a firm need before the share means anything?

Enough that one answer flipping cannot be mistaken for a trend. Share of answers is a count divided by a count, so the smallest reportable movement is fixed by the size of the sample before any engine is involved. The arithmetic below is not an error margin and it is not a confidence interval. It is the floor: the smallest change the instrument can register at all.

The practical read is that a single-engine, single-run check cannot support a month-over-month claim, and a sixty-answer market sample can. Citorian’s 375 runs covered five metros at once, so a single firm measuring one market is working at roughly a fifth of that scale and should plan the prompt count accordingly.

Illustrative arithmetic for a single market. The third column is the smallest movement the sample can register, not a margin of error.
Prompts x engines x runsAnswers in the sampleOne flipped answer moves the share by
1 x 1 x 11100 points
5 x 1 x 1520 points
5 x 1 x 3156.7 points
5 x 4 x 3601.7 points
10 x 4 x 31200.8 points

Which engines are worth running every month?

The ones that both name firms and send people. Citorian narrowed its own roster for the Q3 2026 re-run to ChatGPT, Gemini, Google AI Mode and Google AI Overviews, dropping Perplexity and Claude on the stated ground that together they send under a tenth of AI referral traffic to websites, and adding Google AI Mode.

Consumer reach points the same way. The iLawyerMarketing survey of 1,110 US adults published on August 17, 2026 found 71.9% would use Google to research which lawyer to hire and 41.9% would use ChatGPT, against 15.1% for Gemini, 6.2% for Claude and 3.7% for Perplexity. Firm AEO reads that as four surfaces worth a monthly run and two worth a quarterly one, and notes that a firm cutting engines is trading breadth for a bigger sample on the engines that matter.

What should the monthly ledger record?

One row per run, with the sources column treated as the real output. Which firms got named is the headline, but the list of domains the engines cited is the actionable half, because that list is the retrieval pool the firm has to be present and consistent in. Citorian recorded which firms were named, in what order, whether the engine declined to name anyone, and what sources it cited.

  • Run date, engine, engine version or app surface, and whether the session was logged out.
  • The prompt, verbatim, and its type: direct, vetting, specific-matter, question-led or situational.
  • Firms named, in the order the answer named them. Order carries information that a yes-or-no appearance flag throws away.
  • The outcome when no firm was named, split into declined to name anyone, named only directories, and run failed.
  • Every domain the answer cited. Count the domain column monthly; anything appearing three or more times is the corroboration set for that market.

Which of these numbers arrive in a dashboard instead?

Almost none of them. The Bing Webmaster Tools AI Performance report, released in public preview on February 9, 2026, is the only first-party citation dashboard an engine publishes. It reports total citations, average cited pages, page-level citation activity and grounding queries across Copilot, AI summaries in Bing and selected partner integrations. Microsoft states that it carries no click data and that the grounding queries shown are a sample rather than complete citation activity.

Referral analytics answers a different question and answers it badly. Martindale-Avvo reported in April 2026 that AI tools account for under 5% of total law firm website traffic, and Great Jakes reported AI referrals at 0.47% of sessions across its law firm client portfolio. Those figures describe clicks, not citations, and most AI referrals arrive without a usable referrer. Firm AEO treats an intake field asking how the caller found the firm as the attribution instrument and leaves the GA4 channel-grouping work to its own post.

What can a firm say publicly about its own AI visibility number?

Anything it can date, describe and reproduce. ABA Model Rule 7.1 prohibits false or misleading communications about a lawyer’s services, and a share-of-answers figure becomes misleading the moment it is published without the method attached. A firm quoting its own measurement states the prompt set, the engine list, the run count, the sample size and the date, because those five facts are what make the claim verifiable rather than decorative.

Two claims do not survive that test. An engine naming a firm is not a comparison any tribunal has made, so restating it as a ranking or as proof of being the best converts a measurement into an unverifiable superlative. And an appearance rate measured in one month is not a promise about the next one, since the engines re-crawl and re-rank on their own schedule. Firm AEO is a marketing company rather than a law firm, and the firm’s ethics counsel decides what goes on the page.

Frequently asked

How often should a law firm run the measurement?

Monthly, in the same week each month. The engines do not re-crawl fast enough for a weekly run to register anything but sampling noise, and a weekly cadence multiplies the cost of the protocol by four without improving the signal. Firm AEO treats the monthly appearance rate as the reportable number and a single mid-month check as a diagnostic, not a data point.

Do AI visibility tracking tools replace running the prompts by hand?

They replace the clicking, not the design. A tool still needs the prompt set, the engine list and the run count decided by somebody who knows the market, and most tools report presence rather than the source list, which is the half of the output that tells a firm what to fix. Firm AEO uses tooling for volume and keeps the source column by hand.

What counts as a good share of answers for a law firm?

There is no published benchmark, and any figure presented as one is invented. The useful comparison is internal: this month against last month on a frozen prompt set, and the firm against the two competitors that the same answers named most often. SOCi’s 2026 index found ChatGPT recommending only 1.2% of the locations it studied across all industries, which is a reminder that absence is the normal starting condition rather than a failure.

Should the prompt set include the firm’s own name?

Include a small number of branded prompts, but score them separately. An answer to who is this firm measures whether the engine holds an accurate entity record, which is worth knowing and is a different problem from whether the engine volunteers the firm to a stranger. Mixing branded and unbranded prompts in one percentage produces a number that moves for two unrelated reasons.

Does location sharing change the result?

It can, which is why the session controls have to be identical every month. ChatGPT added opt-in location sharing for more precise local answers on March 26, 2026. A logged-out session with location permission denied and the market named in the prompt text is reproducible; a session that sometimes has device location and sometimes does not is not comparable to itself.

Get a baseline you can compare next month against.

Firm AEO builds the frozen prompt set for your market, runs it logged out across the engines that matter, and hands back the ledger with the source list that says what to fix first.