Methodology
We asked ChatGPT to recommend a local business the way a real buyer would, with web search left un-forced, and recorded the full search chain, not just what it cited.
Scale
The study is built to survive the stochasticity of LLM output: a core measured set, plus repeated-measurement and intent-robustness add-ons, plus a higher-volume pass for the parts that need named businesses at scale.
| Measure | Count |
|---|---|
| Core recommendation sessions (GPT-5.5, un-forced) | 200 |
| Industries × states × core prompts | 10 × 10 × 2 |
| Variance re-runs (repeated measurement) | 120 (4 passes over 10 industries × 3 states) |
| Sub-service sessions (intent-robustness) | 30 |
| Higher-volume pass for named-business data | about 1,000 |
| Total web searches logged | 5,933 (about 30 per session, max 58) |
| Reasoning tokens per session (avg) | about 5,059 |
| Citations captured | 1,982 |
| Distinct sources cited | 858 |
| Businesses named per answer (avg) | about 8 |
| SERP rank checks (both directions) | 3,431 plus reverse-direction |
Process
Un-forced, full-chain capture. Most AI-search studies force a single search and read the citation: clean, but unrepresentative. We let the model run its real multi-step process, so the data reflects what a paying customer actually triggers, and we logged every query in the chain, not only the handful that produced a visible citation.
One assistant, two settings. Everything here is ChatGPT. The core study runs it the way a buyer gets it today: GPT-5.5 with reasoning and live web search. One part of the analysis needs many more named businesses than deep reasoning runs cheaply produce, namely the rank tests (Finding 1) and the count of distinct winners per market (Finding 5.2), so for that part we ran a higher-throughput ChatGPT pass (gpt-5-chat-latest) to gather business names at volume. Every behavioral finding (the audit, the credential checks, the citation and source mix) comes from the reasoning runs. We point out which is which where it matters.
Date (for reproducibility). Runs June 16–17, 2026, via the Responses API web_search tool, reasoning effort high.
The prompts, verbatim ({X} = trade, {sub-service} = a specific job, {city, state} = location):
- "I'm looking for a reputable {X} in {city, state}. Help me find a good one I can hire."
- "Who are the best, most reputable {X}s in {city, state} that I can hire?"
- "I need {sub-service} in {city, state}. Who should I hire?" (a specific job with no reputability cue: car accident lawyer, water heater repair, dental implants, mortgage refinance, couples therapist, Botox, managed cybersecurity, tree removal, homeowners insurance, kitchen remodel), to test whether the audit depends on us asking for "reputable."
Variance-tested. LLM output is stochastic, so we re-ran cells multiple times and report the key behaviors as ranges, not point estimates.
Scope
ChatGPT specifically, on locally-hired service businesses. Other assistants and other business types (SaaS, ecommerce, retail) were not tested. See Limitations.
Finding 1: Search rankings are decoupled from recommendations
Ranking is neither necessary nor sufficient. 90% of recommended businesses do not rank top-20, and 70% of top-ranked businesses are not recommended. This is the most misunderstood thing about AI recommendation, so it comes first.
1.1 The mechanism: it never runs the query a ranking would answer
ChatGPT never runs the buyer keyword, so your rank for it cannot be a direct input. Its searches average 8.4 words and go after named credibility sources, never the short query a buyer types. What it runs are shortlist and verification queries.
Take "plumber in Dallas." A buyer types that into Google and a ranking decides what they see. ChatGPT does not type it. Its real logged queries for the same job look like the examples in Finding 6.3: a ranking lookup like Dallas plumbers Best Pick Reports top rated, a trust check like site:bbb.org/us/tx/dallas plumber accredited A+, then a license check like site:tsbpe.texas.gov "[company] Plumbing" responsible master plumber. It has already chosen candidates from the authority tier and is confirming them. Your rank for "plumber in Dallas" is never read.
(The mechanism here is observed on the reasoning runs. The rank distributions in 1.2 and 1.3 use the higher-volume pass to supply enough named businesses to check against. See Methodology.)
1.2 Recommended businesses do not rank, and the few that do sit on page 3
80% of recommended businesses rank nowhere in the top 50 of either Google or Bing.
We took the businesses ChatGPT recommended and looked up where each one actually ranks for the head buyer keyword a customer would type ("personal injury lawyer Houston," "plumber Dallas," "managed IT services Chicago," "tree service Phoenix"). The answer, overwhelmingly, is nowhere:
| Organic position | Bing | |
|---|---|---|
| Top 3 | 0% | 0.3% |
| 4–10 | 4.1% | 0.9% |
| 11–20 | 5.5% | 0% |
| 21–50 | 10.3% | 0% |
| Nowhere (>50) | 80.1% | 98.7% |
Only 5% rank top-10 in either engine. The minority that rank average position 23, page three.
1.3 And ranking is not sufficient either (the reverse direction)
Of businesses that rank Google top-10 for the buyer keyword, only 30% get recommended.
We ran it the other way too. For each of those same buyer keywords ("dentist Phoenix," "mortgage broker Dallas," and so on), we took the businesses Google ranks in its top 10 and checked how often ChatGPT recommends them back:
| Industry | Top-10-ranked → recommended |
|---|---|
| Medical spa | 39.8% |
| Therapist | 36.8% |
| Home renovation | 36.5% |
| Plumber | 36.2% |
| Insurance | 32.9% |
| Mortgage | 28.1% |
| Arborist | 27.7% |
| IT services | 26.0% |
| Lawyer | 21.7% |
| Dentist | 17.2% |
| Overall | 29.7% |
And within that top 10, rank position barely moves the needle. Pooling all ranked businesses by position, the top 3 are recommended about 26% of the time and positions 4 to 10 about 24%, essentially flat. Being ranked first (31%) is no better than eighth (30%). So it is not even true that ranking higher helps once you are on page one.
So 70% of top-10-ranked businesses are not recommended. That 30% is higher than chance alone would yield (the model names only a handful of businesses out of the many operating in each market) but nowhere near the near-100% you would expect if rank drove recommendation. The most economical read: rank and recommendation share an upstream cause (a prominent business tends to both rank and land in the authority rankings), rather than rank being an input. Ranking is not a direct input to recommendation, and at best weakly correlated through prominence.
1.4 "Search rank" is not even one thing: Google and Bing barely agree
Of businesses that rank top-50 on at least one engine, only 4% rank on both Google and Bing.
| Where a top-50-ranked business appears | Share |
|---|---|
| On both Google and Bing | 4% |
| On one engine only | 96% |
Any "ChatGPT runs on Bing, so optimize Bing" assumption collapses here. The two engines surface almost totally different businesses for the same query. Both engines were pulled through the same SERP API under matched conditions (identical location, language, device and depth, de-personalized), so this divergence reflects genuine engine differences, not session or geo personalization. There is no single "ranking" to win, which is consistent with the model not relying on either engine's results page to choose who to recommend.
1.5 The contrast with citation studies
Search rank decides citations, not recommendations. Ranking and recommendation overlap on 30% of cases at most, and only 5% of recommended businesses rank in any top 10. Citation-focused research (such as the widely-cited query fan-out work) finds that retrieval rank strongly predicts whether a page gets cited in an informational answer. Read that carefully: the "rank" in that work is a page's position inside ChatGPT's own retrieved results, not its Google or Bing ranking. This study measures traditional Google and Bing organic rank, the thing businesses actually optimize. They are two different ladders, and that research is not in conflict with this study, because it measures a different outcome. Our data is about recommendation, and there Google and Bing rank does not predict who gets named. Measured in both directions, the overlap is weak:
| Overlap between ranking and recommendation | Rate |
|---|---|
| Top-10 ranked → recommended | 30% |
| Recommended → ranks anywhere in top 50 | 20% |
| Recommended → ranks in top 10 | 5% |
Traditional rankings and recommendation are close to decoupled. Rank can help your content get quoted. It does not get your business handed to a buyer.
Key takeaways (Finding 1): the model never runs the buyer keyword, recommended businesses overwhelmingly do not rank, ranked businesses mostly are not recommended, and "search rank" fragments across engines (4% Google and Bing overlap). Rank is neither necessary nor sufficient.
Finding 2: Your own site gets you into the running, not recommended
About half of everything ChatGPT cites in the real audit is a business's own site (roughly 52% of the 1,982 citations, classifier-estimated). Here the data corrected our own first guess: we expected business websites to be marginal in the recommendation path, and they are not. The model genuinely reads your pages. But what it reads them for, and which pages, is the whole story.
2.1 Your site is cited about half the time, but the rate swings wildly by industry
Your own site is about half of every citation the model makes, but how heavily it leans on your pages swings from 10% (plumbers) to 74% (med spas) by industry.
| Industry | Citations that are the business's own site |
|---|---|
| Medical spa | 74% |
| Therapist | 73% |
| Arborist | 70% |
| Lawyer | 68% |
| IT services | 54% |
| Dentist | 54% |
| Insurance | 44% |
| Mortgage | 34% |
| Home renovation | 23% |
| Plumber | 10% |
| Overall | about 52% |
The pattern: the model leans on your own site where third-party records are thin (therapists, med spas and arborists have few directories, so it reads the businesses' sites directly) and barely touches it where strong third-party records exist (a plumber recommendation is about 90% third-party: BBB, contractor boards, BuildZoom). Your own site is a fallback the model uses when it has nothing better, not the deciding source. (The business-versus-third-party split is classifier-estimated, so read these as directional.)
2.2 And it is mostly your homepage, used to confirm you exist and offer the service
When the model reads your site it is mostly your homepage: of the business pages we have full URLs for, about 70% were the homepage and 30% a deeper service or location page.
| Industry | Homepage (/) | Deeper page (service or location) |
|---|---|---|
| Insurance | 80% | 20% |
| Home renovation | 78% | 22% |
| Arborist | 74% | 26% |
| Lawyer | 72% | 28% |
| Mortgage | 72% | 28% |
| Medical spa | 70% | 30% |
| Dentist | 69% | 31% |
| IT services | 68% | 32% |
| Therapist | 57% | 43% |
| Plumber | 55% | 45% |
| Overall | about 70% | about 30% |
The deeper pages are exactly what you would expect: service and location landing pages like rotorooter.com/newyork/, nycarborist.com/manhattan, blondiestreehouse.com/arborist-services/ and emeraldtreecare.com/about/meet-our-arborists/. Trades that organize by service or city (plumbers and therapists) get their deeper pages read most often, up to 45% of the time. The buyer's intent shifts it too: a specific, urgent question ("my water heater just burst, who do I call?") pulled deeper service pages 39% of the time, versus 15% for a generic "I need a plumber." The more specific the job, the more the model reaches past the homepage. But even then the homepage leads. Service and location pages do matter, they are how the model confirms you serve this job in this place, but the homepage is still the primary check.
2.3 But never on its own: every recommendation was credibility-checked
Your site alone is never enough: every recommendation was also credibility-checked. Across the 196 sessions that produced a recommendation, all 196 ran a credibility check, zero were standalone, and each came with about 17 verification searches. We checked every session for a credibility check (a verification search, or a citation to a third-party credibility source) in the same session it read the business's pages. It always did.
| Across the 196 sessions that produced a recommendation | Result |
|---|---|
| Recommendations that also ran a credibility check | 196 of 196 (100%) |
| Standalone recommendations (named on own pages, no credibility check) | 0 |
| Credibility-verification searches per recommendation (avg) | about 17 |
Not one business was recommended on the strength of its own website alone. The homepage that gets cited 70% of the time is read in the same session as an average of about 17 credential and trust verification searches. Your pages confirm you are real and relevant, they never substitute for the audit. No single business site recurs across recommendations the way the BBB or a state board does (Finding 5.1): your pages confirm you, they do not differentiate you. Getting your site right makes you eligible. It comes nowhere near getting you recommended.
Key takeaways (Finding 2): your own site is about half of all citations and the model really does read it, mostly your homepage, but never alone. Every recommendation was also credibility-checked (0 standalone, about 17 verification searches each). Your pages make you eligible. Third-party credibility is what gets you recommended.
Finding 3: Reddit and user-generated content are not used
Reddit was cited zero times in the 1,982-citation audit. "Get big on Reddit" has become almost as common a piece of AI-visibility advice as "get more brand mentions," and in the mode that actually recommends businesses, it does nothing.
3.1 Zero Reddit citations in the real audit
Reddit was cited zero times in the real audit, out of 1,982 citations. In a forced single-search pass it appears 171 times.
| Pass | Reddit citations | Total citations | Reddit share |
|---|---|---|---|
| Real reasoning audit | 0 | 1,982 | 0% |
| Forced single-search pass | 171 | 13,655 | 1.3% |
In the reasoning audit, Reddit was cited zero times. The only user-generated or social citations of any kind were two LinkedIn business profiles. Quora, Facebook, forums: zero.
3.2 Reddit only shows up when the model searches shallow
Reddit shows up only when the model searches shallow: 171 citations in the forced single-search pass, against 0 in the real audit. A one-shot search lands on whatever ranks for the query, and Reddit threads rank well, so they get pulled in. The real reasoning audit never does that. It issues targeted queries against named credibility sources (boards, associations, the BBB), and a Reddit thread is none of those. So Reddit citations are an artifact of shallow retrieval, not a signal the model uses to recommend. This matters because naive checks and many AI-visibility tools rely on exactly that kind of one-shot call, which is why they over-report Reddit's influence.
Key takeaways (Finding 3): the recommendation audit cited Reddit zero times. It appears (171 times) only in a shallow single search, which is not how recommendations are made. Forum presence does not move a recommendation.
Finding 4: The platforms you pay for are bystanders
The five platforms businesses most often pay for total just 8 of 1,982 citations. The directories you budget for barely register in the audit at all.
4.1 Paid platforms in the real audit
The big five paid platforms (Yelp, Angi, Avvo, Thumbtack, HomeAdvisor) total just 8 of 1,982 citations.
| Platform | Citations (of 1,982) |
|---|---|
| Healthgrades | 9 |
| Yelp | 6 |
| Angi | 2 |
| Trustpilot | 2 |
| Thumbtack / HomeAdvisor / Porch | 0 |
| Avvo / Justia / FindLaw | 0 |
| Nextdoor / Bark / Networx | 0 |
Healthgrades, a medical directory, is the single most-cited paid platform at 9, and still negligible. Vitals, Martindale and Lawyers.com are at or near zero.
4.2 They show up shallow, like Reddit
Like Reddit, paid platforms are a shallow-search artifact: HomeAdvisor goes from 0 in the real audit to 22 in the forced pass, Yelp from 6 to 12.
| Platform | Real audit | Forced single-search pass |
|---|---|---|
| Healthgrades | 9 | 26 |
| HomeAdvisor | 0 | 22 |
| Yelp | 6 | 12 |
| Angi | 2 | 6 |
| Avvo | 0 | 3 |
Paid directories appear several times more often in the shallow forced pass, where the model grabs whatever ranks, than in the real audit, where it verifies credibility. These are advertising marketplaces, and the audit is not looking for ads. Paying to appear on a platform the audit never consults is spend that does not touch the recommendation (it may still drive direct traffic, a separate question).
Key takeaways (Finding 4): the paid directories businesses budget for are almost entirely absent from the audit (8 of 1,982 for the big five), and what little shows up is a shallow-search artifact. Ad spend on these platforms does not buy a recommendation.
Finding 5: There is no universal playbook
ChatGPT cited 858 distinct sources. 95% appear in only one industry, and 68% were cited only once. The decisive sources are specific, fragmented and unoptimized.
5.1 The sources are radically industry-specific, and overwhelmingly one-off
68% of all cited sources appeared exactly once, and the average source spans just 1.1 of 10 industries. The only sources that cross industries are generic trust and government, not anything you would "market into":
| Cross-industry source | Industries (of 10) | Type |
|---|---|---|
| mass.gov | 8 | govt umbrella portal |
| bbb.org | 7 | trust directory |
| pa.gov | 6 | govt umbrella portal |
| content.boston.gov | 5 | govt portal |
| expertise.com | 5 | listicle |
| (everything else) | 4 or fewer. 95% appear in exactly 1, 68% cited only once | industry-specific |
The top 20 sources are only 34% of all citations, a long fragmented tail. This does not contradict the audit's "same stack everywhere" (Finding 6). The structure is universal: discover, verify, trust, the same in every market. What fills each tier is not. The exact sources vary by industry, by state, and even by service within an industry. The shape is fixed, the contents are local. That is why there is no list to copy, only a process to run for each market.
5.2 The winners are hyper-local too: it is 100 separate local races
Pool about 100 recommendations per industry and you get 203–425 distinct businesses. No national winner.
We finally looked at who gets recommended, not just the sources, and the winners are as fragmented as the sources. This winner count uses the higher-volume pass (about 100 recommendations per industry), because counting distinct winners needs many more named businesses than the deep reasoning runs produce.
| Signal | Result |
|---|---|
| Distinct businesses recommended per industry (about 100 recommendations) | 203–425 |
| Most-recommended business's share of an industry | 20% at most (no dominant national winner) |
| Within a single city, how often the leading business recurs across differently-worded queries | 78% (sticky per market) |
| Exception, a national brand winning across markets | Roto-Rooter (plumbers), 20 of 100 |
The reconciliation: AI recommendation is about 100 separate local races. Within one market the leader is locked in (78% consensus across phrasings), but across markets the winners are all different, hundreds of them, because the audit resolves locally. National brands can occasionally win broadly (Roto-Rooter), but most verticals have purely local champions. So it is not just the sources that resist a generic playbook. The winners do too.
5.3 BBB is the one near-universal trust anchor
A single source, the BBB, is 13% of every citation in the study (261 of 1,982). It dominates where a profession has no prestige system of its own, and disappears where one exists:
| Industry | BBB citations | Has its own prestige ranking? |
|---|---|---|
| Plumber | 81 | no |
| Insurance | 62 | no |
| Mortgage | 56 | no |
| Home renovation | 51 | no |
| Arborist | 8 | no |
| IT services | 2 | yes |
| Lawyer / Dentist / Therapist | 0 | yes |
Checked by exact profile URL, by name (site:bbb.org/us/tx/houston/profile/plumber … A+). Profession-specific authority where it exists, BBB where it does not.
5.4 Listicles, editorial press and media coverage: a top lever in some industries, irrelevant in others
Editorial press (city magazines, "best of" listicles, trade press) is 13.5% of citations overall, but that average hides a split. It appears in 85% of lawyer and IT sessions and 0% of therapist and arborist sessions.
| Industry | Sessions with an editorial source (of 20) | Dominant editorial sources |
|---|---|---|
| Lawyer | 17 | Best Law Firms, Chambers, Best Lawyers, Forbes |
| IT services | 17 | Clutch, CRN, Channel Futures |
| Dentist | 13 | Seattle Met, Philly Mag, Houstonia, Castle Connolly |
| Plumber | 5 | Best Pick Reports, Expertise |
| Medical spa | 4 | city magazines, RealSelf |
| Mortgage | 4 | Expertise, WalletHub |
| Insurance | 2 | n/a |
| Home renovation | 2 | NY Magazine, Houzz |
| Therapist | 0 | (runs on associations instead) |
| Arborist | 0 | (runs on certifications instead) |
Why the split. Editorial press is how ChatGPT fills the authority-ranking tier (tier two). Fields with a strong third-party ranking culture give it plenty to cite: law has Best Law Firms and Chambers, IT consulting has Clutch and CRN, dentistry has city "Top Dentists" features. Fields that run on membership and certification instead (therapy on APA and Gottman, arboriculture on ISA and TCIA) give it nothing of that kind to cite, so it falls back to those directories. Unregulated IT shows a second effect on top: with no licensing board to check, listicles like Clutch become the credibility proxy the license tier would otherwise provide. So it is not simply "no regulator means more press." Law has a strong regulator and heavy press. The driver is whether the field has a public ranking and press culture at all.
The authority tier is two distinct things, which in SEO are not interchangeable: review aggregators (Clutch, RealSelf, Zillow, which rank businesses by collected ratings and reviews) and editorial "best of" rankings and listicles (Best Law Firms, Chambers, Expertise.com, a city magazine's "Top Dentists", where a publication curates the list on its own methodology). Which one the model leans on depends entirely on the field:
| Source | Type | Citations | Where it carries |
|---|---|---|---|
| Clutch | review aggregator | 67 | IT services (almost only) |
| Best Law Firms | editorial ranking | 66 | Law |
| City magazines ("Top X") | editorial listicle | 47 | Dentist, med spa |
| Expertise.com | editorial listicle | 18 | Plumber, mortgage |
| Best Lawyers | editorial ranking | 15 | Law |
| Chambers | editorial ranking | 12 | Law |
| Forbes | editorial media | 8 | Law |
| RealSelf / Zillow | review aggregator | 1 each | med spa, mortgage |
Editorial rankings and listicles carry the bulk of the citations (about 170), review aggregators far fewer (about 69, almost all Clutch). And aggregators own exactly one field here: IT services, where Clutch is the de facto category ranking. Everywhere else, editorial rankings and listicles dominate. So the play is vertical-specific: chase the aggregator if you are in IT (Clutch), and chase the editorial ranking or listicle otherwise (Best Law Firms, a city "Top X" feature). A placement in the right one for your field is a direct recommendation lever, and close to wasted effort in therapy or tree care.
5.5 Professional associations are the biggest hidden layer
Associations are the dominant authority tier for certification- and membership-driven fields like arborists, therapists and remodelers, and almost no one markets to them. For arborists the association directories (treesaregood and TCIA, 21 citations) outrank the firms themselves.
| Association (citations) | Industry it carries |
|---|---|
| NARI: nariatlanta (18), greaterphoenixnari (9), nari.org | Home renovation |
| ISA / TCIA: treesaregood (11), treecareindustryassociation (10) | Arborist (top sources, above the firms) |
| TrustedChoice / Big "I" (8) | Insurance |
| ADA: findadentist (4) | Dentist |
| Gottman (3), ABCT (3), APA locator | Therapist |
Recognized membership is a verifiable credibility signal, and the association directory is treated as a trusted shortlist. Active, listed membership is high-leverage and invisible to anyone optimizing Google or Yelp.
Key takeaways (Finding 5): sources and winners are both fragmented (95% single-industry sources, 68% one-off, 203–425 distinct winners per vertical). BBB is the one near-universal lever, editorial press is a major but uneven PR lever, and associations are the biggest hidden layer. The stack is different for every industry and state.
Finding 6: How it really works, the credibility audit
Every recommendation is a roughly 30-search investigation. This is the credibility layer in full, what all of the above points to. When a buyer asks ChatGPT to recommend a provider, it does not read a rankings page and copy the top results. It runs a credibility audit: across 200 sessions it averaged around 30 web searches and about 5,059 tokens of reasoning per recommendation, consulting roughly seven distinct sources in a consistent order: discover candidates, verify their credentials, cross-check a trust directory, then name eight or so businesses. The shape was identical in every industry.
6.1 Recommendation is a 30-step search process
It runs about 30 searches per recommendation, and two-thirds of them are hidden verification that never becomes a visible citation.
| Metric | Value |
|---|---|
| Searches per session | avg about 30, max 58 |
| Internal reasoning per session | about 5,059 tokens |
| Distinct sources consulted per session | about 7 |
| Businesses named in the final answer | avg 8.2, max 13 |
| Searches per citation produced | about 3 to 1 |
Roughly two-thirds of the searches, nearly all credential verification, never surface as citations, and about half of what it does cite is a regulator, ranking or directory rather than a business (about 50%, classifier-estimated). There is a mild ordering tendency: discovery and ranking queries skew slightly earlier in the chain than credential and trust queries (mean normalized position 0.41 versus 0.49). It is a small gap, so read it as a lean toward "discover first," not a clean two-phase sequence. Either way, the search chain, not the citation list, is the unit of truth.
6.2 Every industry runs the same three-tier stack
Same three tiers in all 10 industries. Only the contents change.
| Industry | ① Licensing / credential | ② Authority ranking | ③ Trust directory |
|---|---|---|---|
| Lawyer | State Bars (calbar, texasbar…) | Best Law Firms, Chambers, Super Lawyers | bar referral |
| Plumber | Contractor boards (CSLB, AZ ROC, WA L&I) + permits | Expertise, Best Pick Reports | BBB |
| Dentist | Dental boards + ADA find-a-dentist | City "Top Dentists" magazines, Castle Connolly | Zocdoc |
| Mortgage broker | NMLS + CFPB + state DFI/DRE/DFS | Expertise, Zillow | BBB |
| Therapist | State licensing + APA/ADAA locators | n/a | Psychology Today |
| Medical spa | FDA + state medical/cosmetology boards | RealSelf, city magazines | BBB |
| IT services | n/a (unregulated) | Clutch, CRN, Channel Futures | Microsoft Partner |
| Arborist | ISA / TCIA / ASCA certification | n/a | BBB |
| Insurance agency | State Insurance Depts (CA DOI, TX TDI) | n/a | BBB, TrustedChoice |
| Home renovation | Contractor boards + NARI | Houzz, NY Magazine "best contractors" | BBB, BuildZoom |
The tiers adapt. Unregulated IT drops the licensing tier. Therapy and insurance, with no prestige ranking, lean on associations and trust directories instead.
6.3 The model's queries are nothing a human would type
Its searches average 8.4 words (a human's is about 4), and 18% use a site: operator. These are machine-shaped queries no human would type.
| Query trait | Value |
|---|---|
| Average length | 8.4 words |
| Longer than 6 words | 77% |
Uses a site: operator | 18% |
| Verifies a name in quotes | 14% |
| Carries a recency year (2023–2026) | 9% |
| Broad "brand buzz" sweeps | 0 |
Verbatim examples:
site:mqa-internet.doh.state.fl.us "Leslie Haller" "Dentist" "Coral Gables"
Abraham Watkins Houston personal injury attorneys board certified Best Law Firms Tier 1
https://data.wa.gov/resource/m8qx-ubtq.json?$select=businessname (a raw state open-data API endpoint)
To verify a Washington contractor, ChatGPT queried a government open-data API directly. These queries are synthesized on the fly and effectively infinite in variety. You cannot optimize for a query you cannot see, and you cannot see these without capturing them from real ChatGPT calls.
6.4 The audit holds without asking for "reputable"
Strip the "reputable/best" cue and credential verification still fired in 93% of sub-service sessions.
| Sub-service query | Credential-verified | What it specialized to |
|---|---|---|
| dental implants | ✓ | TX dental board, UCLA dentistry, Mount Sinai |
| Botox | ✓ | Mount Sinai, Weill Cornell, CDC |
| couples therapist | ✓ | Gottman Referral Network |
| mortgage refinance | ✓ | CFPB, Freddie Mac |
| tree removal | ✓ | TCIA, city forestry |
| homeowners insurance | ✓ | NY DFS, CA DOI, TX TDI |
| car accident lawyer | ✓ | State Bar, courts |
| water heater repair | ✓ | city permits, BBB |
93% of sub-service sessions (28 of 30) verified credentials, at 24.4 searches each. The audit is intent-robust, not an artifact of our wording.
6.5 It verifies against the specific licensing body for each state
Credential verification is the dominant behavior across the study, firing in 74–96% of regulated-industry sessions. It vets about six candidates by name per session, and checked a license by name in 40% of sessions.
| Measure | Result |
|---|---|
| Sessions touching a government/regulator source | 90% (98% of regulated industries) |
Named-individual license checks (site:[board] + a name) | 40% |
| Distinct candidates verified by name per session | avg 5.5, max 51 |
| Regulated-industry verification (variance-tested) | 74–96% (mean 87%) |
| Regulated markets verifying at least once | 100% |
The credential check is the dominant, recurring behavior. Every regulated market triggers it, but it is not 100% every time. On a given run, roughly 1 in 5 regulated recommendations skip the explicit license step. The behavior is robust, the rate is stochastic, which is why we report a range. For health verticals it also reaches academic and clinical institutions (about 39 .edu citations: UCLA/USC dentistry, Mount Sinai for med-spa). What we measured is a query pattern. Lawyers route to the State Bar, mortgage brokers to NMLS, contractors to the state registry. The functional claim is supported. The cognitive claim ("it knows") is interpretation.
6.6 The right board appears for each state
Each industry surfaced 16–31 distinct regulators, and the correct state's board showed up in most runs.
| Add this state… | …and the model verifies against |
|---|---|
| Arizona | azbar.org, AZ Registrar of Contractors, azdentalboard.us |
| Florida | floridabar.org, myfloridalicense.com, FL Dept. of Health |
| Washington | wsba.org, secure.lni.wa.gov, doh.wa.gov |
| Texas | texasbar.com, Texas Board of Legal Specialization, tsbpe.texas.gov |
In some states one portal serves many trades. mass.gov was the verification source across 8 different industries.
6.7 Volume loses to precision
In 5,933 searches, ChatGPT ran zero broad "what's-being-said-about-this-brand" sweeps. Every search is a targeted query against a named source. And when it does query a specific business by name, it never does so nakedly: every branded query in the sample paired the name with a verifier, a site: check against a licensing board, a BBB profile lookup, or a qualifier like "Best Law Firms Tier 1" or "complaints."
| Query behavior | Of 5,933 searches |
|---|---|
| Broad brand-buzz sweeps ("what's being said about X") | 0 |
| Bare-brand exploration (a name with no verifier) | 0 |
| Names a specific business, always with a verifier attached | 14% |
So "get more brand mentions everywhere" misses. A mention on fifty mid-authority sites is on zero of the sources the audit queries, while the one right source (your board, association or category ranking) is checked directly. The shotgun's only mechanism is seeping into a future model's training data, a bet that pays off at the next model cycle (months to years out), not in front of today's buyer, who is served by a live audit that mentions do not touch.
Key takeaways (Finding 6): a 30-search investigation per recommendation, two-thirds of it hidden verification, structured as a universal three-tier stack, verifying against the specific state licensing body, with precise machine-shaped queries that never explore a naked brand name.
How to actually get recommended
This is the part that matters. Everything above says what does not move a recommendation. Here is what does, in the order the audit cares about. Because the deciding sources are specific, fragmented and per-market, this is not a list to copy, it is a process to run for each market you serve.
- Cover the basics so you are in the running. ChatGPT reads your own pages to confirm you exist, offer the service, and operate in the market, and that is about half of what it cites (Finding 2). A homepage, a page per service, and a page per location clear this bar. Most sites already have it. This does not get you recommended, it makes you eligible.
- Win the authority tier. The highest-leverage work, and it is per-market: a blend of review aggregators where they carry weight (Clutch for IT services), curated "best of" listicles (Expertise, a city magazine's "Top X"), and the specific editorial rankings worth going after in your field (Best Law Firms, Chambers). Capture which of these the model actually pulls for your category and city (Finding 5.4), then earn placement on them.
- Be in good standing on the credential tier. The board changes by industry and state (Finding 6.6). Be correctly licensed, listed and clean on the exact regulator the model checks for your market, by name.
- Confirm the trust tier. BBB by exact profile URL in trades and financial services, the relevant association directory in the professions (Findings 5.3 and 5.5). Make that profile accurate, accredited and well-rated.
- Repeat for every service and every location. Every service and every city resolves to a different source list and a different set of winners (Findings 4 and 5). What gets a Dallas plumber recommended is not what gets a Houston one recommended, and tree removal is not tree trimming. Each is its own job.
- Re-run on a cadence. The behavior is stochastic and drifts with model versions. Re-capture periodically to catch new opportunities, hold your place on the shortlist, and stay current as newer versions of ChatGPT ship.
Why you cannot do this by hand
A single ChatGPT answer is not evidence. The output is stochastic, so the same prompt run twice gives different names, different sources and a different search chain. We measured this directly: regulated-industry verification swings between 74% and 96% across re-runs, roughly 1 in 5 regulated recommendations skip the license check on any given pass, and within a single market the leading business only recurs 78% of the time across reworded queries. To get a stable read on even one market you have to sample the same question many times, across phrasings, across the surrounding sub-services, and re-sample as the model updates. That is thousands of calls per market. Eyeballing a few answers tells you nothing reliable, and it is exactly the trap most "I checked ChatGPT and we're not there" reactions fall into.
Why it is now within reach for anyone
The good news is that the raw capability is not gated. You do not need a premium AI-visibility tracking suite. The whole method is the model's own API plus a capable general AI (Claude, for instance) to drive the sampling, parse the full search chain out of each response, and synthesize the patterns into the per-market source list. The API returns the queries and citations. The AI does the orchestration and the reading. Everything in this report was produced that way. The information is observable by anyone willing to capture it properly, which is the entire point: the audit is hidden, but it is not secret.
Limitations
- Local-service businesses only. All ten industries are locally-hired services. SaaS, ecommerce, retail and other categories were not tested. The machinery should extend with different tiers (G2 and Capterra and security review for SaaS, marketplace ratings for ecommerce), but that is untested here. The unregulated IT vertical, which dropped the licensing tier for Clutch and CRN, is an early hint.
- One named engine, dated. ChatGPT (GPT-5.5), June 2026, not "AI" in general. Other assistants weight sources differently and behavior shifts across versions. ChatGPT is the dominant consumer assistant, on the order of roughly 900M weekly users, which is why it is the right single focus.
- Stochastic, and moving. Key rates are reported as ranges to account for run-to-run variation, and we did not measure day-to-day index drift over weeks. Both the model and the specific businesses it cites will keep changing as ChatGPT updates, so treat every named source and business here as a snapshot, not a constant. The method is the durable part, not any single name.
- Two figures are automated estimates, not hand-counts, so treat them as approximate: the roughly 50/50 split between credibility sources and businesses in the citations, and the distinct-winner tallies in Finding 5.2 (which may include a few directory pages that look like business sites).
Scope: 200 reasoning sessions on GPT-5.5 (plus sub-service and variance runs), plus a higher-volume ChatGPT pass for named-business, rank and winner data. 10 local-service industries, 10 states, 5,933 logged searches, 1,982 citations, 3,431 plus reverse-direction rank checks. ChatGPT, June 16–17 2026.