We Asked ChatGPT to Recommend a Lawyer 15 Times. The Same Names Kept Winning.
We ran 15 recommendation prompts across five US cities. Every shortlist came from the same third-party directories — and not one cited a law firm's own website.
We asked ChatGPT to recommend a personal injury firm 15 times — five US cities, three separate runs each. The shortlists barely moved. Ten firms were named in all three runs of their city, every answer built its list from third-party directories rather than firms' own websites, and only one national brand crossed city lines.
Here is the full result, the method behind it, and what it means if you are trying to get recommended by an AI answer engine.
What We Ran
Five metropolitan markets. Three independent runs per city, each in a fresh session with no prior context and no brand names supplied by us. Fifteen prompts total, logged in July 2026.
We deliberately kept the prompt blunt — the kind of question a buyer with no shortlist actually asks — and we recorded three things each time: which firms were named, in what order, and which sources the answer drew on.
Finding 1: The Same Names Kept Coming Back
Ten firms appeared in all three runs for their market. The repeat rate across runs sat in the region of 60 to 70 percent, which means the shortlist is far more stable than the "AI answers are random" narrative suggests.
That stability cuts both ways. If you are on the list, you are on it durably. If you are not, you are not going to fall onto it by accident.
Only one firm — Morgan & Morgan, a national brand — appeared across multiple cities. Every other winner was a local incumbent. AI recommendation, at least in this category, is a local competition being decided by national-scale sources.
Finding 2: Not One Answer Cited a Firm's Own Website
This is the finding that should change how you spend your budget.
Every single shortlist was assembled from third-party sources: Forbes, Best Lawyers, Expertise, SuperLawyers, and Chambers. Across 15 runs, no answer sourced its recommendation from the website of the firm it was recommending.
Read that again if you are about to commission a homepage redesign. The model was not reading these firms' sites to decide who was good. It was reading what other people had published about them, and the firms' own marketing had no vote.
Finding 3: The Winners Shared Three Traits
Looking at what the consistently-named firms had in common, three things separated them — and none of them is a copywriting exercise.
Directory dominance. The winners were present, complete, and well-positioned on exactly the directories the model was citing. Not one directory. The stack of them.
Review depth, not just review scores. A high rating on thin volume did nothing. The firms that surfaced repeatedly had both — Hirsch & Lyon, for example, carried 282 Google reviews at a 4.9 average. Volume at a high average reads as evidence; a 5.0 on eleven reviews does not.
Demonstrable, specific results. The named firms published concrete outcomes, and those outcomes were picked up elsewhere. Arnold & Itkin surfaced with $25 billion recovered and an $8 billion verdict attached to their name. Kramer Dillof carried Inner Circle of Advocates membership and a $172 million verdict. These are facts a model can repeat with confidence, because they are specific, checkable, and published in more than one place.
What This Means If You Are Not a Law Firm
Legal is an unusually directory-dense category, so do not port the specifics across without checking. The mechanism, though, generalises cleanly.
Every category has a set of sources that AI answers lean on. For SaaS it might be review platforms and comparison sites. For local services, maps and review aggregators. For B2B, industry press and analyst coverage. The question is not whether your category has this layer — it is whether you have identified it and whether you are present in it.
The generalisable lesson is this: your website is where you get cited, but third-party sources are where you get considered. If you are absent from the second, the first never gets a chance.
The Uncomfortable Implication
Most brands invest in the surface they control and neglect the surfaces they do not. That was defensible when a website was the thing being ranked. It is much harder to defend when the deciding layer sits entirely off your domain.
None of this means your site does not matter. When a model does reach your pages, structure and substance determine whether you get quoted or skipped — which is a real, separate discipline covered in our guide to AI search visibility. But being reachable is downstream of being considered, and consideration is happening somewhere you probably are not measuring.
What We Would Do With This
If we were starting from these findings on a new engagement, in order:
- Identify the citation set. Run the category prompt several times and write down every source the answers name. That list is your actual target, and it is category-specific.
- Audit presence across that set. Completeness and accuracy first, positioning second. An incomplete profile on a cited directory is a wasted slot.
- Build review depth deliberately. Volume at a high average, sustained over time. This is slow and there is no shortcut worth taking.
- Publish specific, checkable results. Numbers, outcomes, named achievements — the things a model can restate without hedging.
- Re-run the prompts monthly. Checking what AI says about you is a measurement discipline, not a one-time diagnostic.
Methodology and Limits
We are reporting this as what it is: 15 runs, one model, one category, one point in time.
It is not a peer-reviewed study, and we would caution against treating the percentages as precise constants. Model behaviour shifts, source availability changes, and a rerun next quarter would not reproduce these lists exactly. What we are confident in is the shape of the finding — the stability of the shortlist, the total dominance of third-party sources, and the traits the winners shared — because those held consistently across all five markets rather than appearing in one.
We would rather publish the real number of runs and the real limits than round the methodology up into something more impressive. That is also, not coincidentally, the kind of specificity that gets a page cited in the first place.
Want to know where your brand stands in your own category? Start with a free audit.