Growli
Blog

How AI Assistants Decide Which Businesses to Recommend

Zakaria Reziki

By Zakaria Reziki

CEO — Growli · August 14, 2026 · 12 min read

Drafted with AI assistance under editorial standards set by Zakaria Reziki, then published after automated sourcing and quality checks.

AI assistants recommend businesses by running a retrieval pipeline: they rewrite your prompt into searches, pull a set of documents, weight those sources for trust, resolve which real-world business each mention refers to, and then synthesise a short answer naming the entities that survived all four filters. No assistant holds a ranked list of "best plumbers in Leeds" waiting to be read out. The list is assembled at question time, from whatever the model could fetch and trust in the seconds after you hit enter.

That distinction is the whole game. If you think of it as an opinion, you try to persuade the model. If you think of it as a pipeline, you go and fix the specific stage where you drop out — and those stages behave differently on ChatGPT, Gemini, Claude, Perplexity and Copilot, which is why the same prompt produces different names on each.

What follows walks the pipeline end to end, with the documented behaviour of each platform where it exists, and marks plainly which levers a business actually controls at every step.

The four stages between a prompt and a name

A prompt like "who should I use for commercial fit-outs in Manchester?" is not a search query. Before anything is retrieved, the assistant decides whether it needs the web at all, then expands the request into several concrete searches — typically a mix of category terms, location modifiers and intent words like "best", "reviews" or "compare". OpenAI describes ChatGPT search as combining its model with information from the web and returning links to sources, rather than answering purely from training data (OpenAI).

Retrieval returns a candidate set: usually a few dozen URLs, of which only a handful fit in the working context. Those get filtered — some sources are treated as reliable for commercial claims, some are ignored, some are excluded because the publisher blocked the crawler. Anthropic's web search for Claude, for example, returns answers with citations to the pages used (Anthropic), which means the sources are load-bearing and inspectable.

Then comes the step most businesses never think about: entity resolution. The assistant has to decide that "Northgate Interiors", "Northgate Interiors Ltd" and the unnamed contractor described in a trade magazine article are the same company — or that they are not. Finally, synthesis: the model writes two to five names into prose, in an order shaped by how strongly and how consistently each candidate appeared in the retrieved evidence.

Every stage is a filter. You can be crawlable, well-reviewed and genuinely the best option and still lose at entity resolution. Diagnosing which filter removed you is more useful than any general advice about "optimising for AI".

  • placeholder
  • x
  • Stage gates
  • :

Recommendation pipeline

From prompt to named business in four filters

  1. Interpret and fan out

    The assistant decides whether the web is needed, then rewrites the prompt into several concrete searches using buyer vocabulary.

  2. Retrieve

    Dozens of URLs are pulled from that platform's index; only a handful fit into the working context.

  3. Weight for source trust

    Independent pages that compare options outrank self-descriptions; blocked or stale sources drop out.

  4. Resolve entities

    Mentions are collapsed into one real business — or split, which weakens every piece of evidence.

  5. Synthesise

    Two to five names are written into prose, favouring candidates with explicit, quotable qualifiers.

Stage 1: retrieval — fetchable first, findable second

Nothing downstream matters if the fetch fails. Each assistant uses its own crawlers and its own index, and each publishes the user agents involved: OpenAI documents OAI-SearchBot for surfacing sites in ChatGPT search alongside GPTBot for training (OpenAI), and blocking one is not the same as blocking the other. A robots.txt written in 2023 to keep AI training out frequently keeps the search bot out too — a self-inflicted exclusion that no amount of content work will fix.

The index behind the assistant matters just as much. Google states that its AI features in Search draw on the same systems and indexing as Search itself, and that the guidance is the same Search Essentials guidance (Google Search Central). Microsoft Copilot inherits Bing's index, so Bing's webmaster guidelines on crawlability, canonical URLs and content quality are the operative rulebook (Bing Webmaster Guidelines). Perplexity runs its own retrieval and cites sources inline, and has built commercial relationships with publishers through its Publishers Program (Perplexity).

Then there is query mismatch. Assistants search in the vocabulary of the user, not the vocabulary of your marketing. Businesses that describe themselves as offering "integrated workplace solutions" are invisible to a fan-out of "office fit-out contractor Manchester", "commercial interior fit-out companies UK" and "best fit-out firms for offices". Category-plain language on your service pages, in the same words buyers use, is what makes a page retrievable for the queries that actually get run.

What you control here: crawler access, page-level indexability, freshness of key pages, and whether your terminology matches the search phrasing an assistant is likely to generate.

Stage 2: source trust — third-party pages usually pick the shortlist

Once documents are retrieved, they are not equal. Assistants lean heavily on pages that read as independent assessments — directories, review platforms, trade press, comparison articles, association member lists, marketplace profiles — because those are the pages that contain lists of competing options. A vendor's own site is excellent for confirming details and terrible as evidence of superiority; "we are the leading provider" is a claim the model has been trained to discount.

This is why the practical shortlist is often set before your website is ever read. If five roundup articles about your category name eight companies and you are in none of them, the assistant's candidate set is those eight. Your site then gets fetched only to enrich a name that already appeared, or not at all. The asymmetry is uncomfortable but actionable: earning a line in the pages that already rank for "best X in Y" moves you further than another homepage rewrite.

Consistency across sources acts as a second trust signal. When six independent pages describe you the same way — same specialism, same city, same size of client — the model treats that description as settled fact. When they conflict, the model hedges, and hedged entities get dropped from short answers because they cost words to explain.

Recency matters more than most teams expect. Assistants prefer documents with current dates for commercial questions, so a 2019 directory listing with a dead phone number is worse than no listing. Auditing where you appear, and getting stale profiles corrected, is unglamorous work with direct effect on retrieval. We cover the measurement side of this in AI visibility.

  • Pages that build shortlists: category roundups, "best of" lists, comparison and alternatives articles, review platforms, industry directories and association registers.
  • Pages that confirm details: your service pages, pricing page, case studies, about page, and structured data.
  • Pages that damage you: outdated profiles with wrong contact details, abandoned duplicate listings, and pages describing a service line you no longer offer.

Source trust audit

What to fix before rewriting another page

  • Appear in the pages that already rank for your category prompts

    Roundups, comparison articles, directories and association registers form the candidate set an assistant works from.

  • Make independent descriptions of you agree

    Six sources describing the same specialism and location turn a claim into settled fact; conflicting sources cause hedging.

  • Kill or correct stale profiles

    An outdated listing with a dead number or a discontinued service is worse than no listing for commercial questions.

  • Publish verifiable differentiators, not adjectives

    Sector, deal size, certification and turnaround are quotable; "award-winning" is discounted.

Stage 3: entity resolution — being one unambiguous business

The model now has a pile of mentions and has to collapse them into entities. This is where near-identical names, rebrands, multi-location operations and holding-company structures cause silent losses. If evidence for your company is split across three half-formed identities, none of the three looks strong enough to recommend, even though the sum would.

Machine-readable identity is the cheapest fix available. Google's structured data documentation for local businesses specifies the properties that describe a physical business — name, address, phone, opening hours, geo, price range (Google Search Central) — and Schema.org's sameAs property exists precisely to link your page to the canonical references for the same entity, such as an official profile or a knowledge-base page (Schema.org). Filling those in gives every crawler the same unambiguous statement of who you are.

The rest is discipline. Use one legal-plus-trading name pattern everywhere. Keep the same category words in your title tag, your directory profiles and your LinkedIn description. Make sure the address and phone number on third-party listings match your site character for character. If you have multiple locations, give each one its own indexable page with its own structured data rather than a single dropdown-driven page.

Assistants also resolve entities by co-occurrence: the brands, technologies, certifications and client types that appear near your name. If nothing in the retrieved evidence connects you to the specific service in the prompt, you are resolved as a generic company in your industry — retrievable, but never the answer to a specific question.

Stage 4: synthesis — how a shortlist becomes a sentence

At synthesis, the model has perhaps 5,000 to 20,000 words of retrieved text and needs to produce 120. Compression decides everything. Candidates supported by explicit, quotable statements survive; candidates that require inference get cut, because the model has no room to argue.

Position inside the retrieved context has a measurable effect. Liu and colleagues showed that language models use information best when it appears at the beginning or the end of a long input, with accuracy degrading noticeably when the relevant passage sits in the middle (Liu et al., "Lost in the Middle"). Practically, that rewards content structure: a claim stated in the first sentence of a page section, or in a clearly labelled comparison table, is far likelier to survive compression than the same claim in paragraph nine.

The answer's shape also follows the shape of its sources. If the retrieved documents are comparison articles, you get a comparison. If they are review pages, you get sentiment. If a source has already framed a market as "three enterprise options and two budget options", the assistant tends to reproduce that framing and slot names into it. Whoever writes the framing that gets retrieved has more influence than whoever writes the best sales copy.

Finally, the model needs a reason to name you. "Good for regulated industries because of its ISO 27001 certification" is a sentence the assistant can lift and defend. "Award-winning and customer-focused" is not. Publishing the specific, verifiable qualifiers a buyer would filter on — sector, deal size, geography, compliance, integrations, turnaround — supplies the differentiators that make you quotable.

Why the same prompt gives different answers on each assistant

Divergence is structural, not random. Copilot answers from Bing's index; Gemini's AI features run on Google's Search systems and indexing (Google Search Central); Perplexity retrieves through its own stack and cites inline; Claude searches the web and attaches citations to the claims it draws from it (Anthropic); ChatGPT search combines the model with web information and links out (OpenAI). Different indexes contain different documents, refreshed on different schedules. Same question, different evidence.

On top of that sit four more sources of variance. Query expansion differs, so the searches actually run are not the same. Sampling makes generation non-deterministic, so the same assistant can produce a different order twice in a row. Localisation shifts results by inferred country and city. And personalisation — chat history or memory features — can bias one user's answers toward brands they have discussed before.

The operational consequence is that a single spot-check tells you almost nothing. To know whether you are recommended, you need the same set of buyer-intent prompts run repeatedly, across assistants, from the geographies you sell into, with the cited sources captured each time. That is what Growli does: we track mentions, comparisons and recommendations per assistant, show which sources the answers were built from, and turn the gaps into a ranked list of fixes — see how it works.

Diagnosis follows the same four stages. Never retrieved anywhere? Suspect crawler access or query mismatch. Retrieved but never named? Suspect source trust or entity ambiguity. Named but described wrongly? A specific stale source is winning, and it can be identified from the citations.

Gap diagnosis

Reading the symptom back to the failing stage

  1. Never retrieved on any assistant

    Check crawler access per documented user agent, page indexability, and whether your wording matches buyer search terms.

  2. Retrieved but never named

    Source trust or entity ambiguity — you appear in weak sources, or your evidence is split across inconsistent identities.

  3. Named on one assistant only

    An index-coverage problem: the document that supports you exists in one platform's index and not the others.

  4. Named but described wrongly

    A specific outdated source is winning the description; the citations in the answer will tell you which one to fix.

What you can actually influence, stage by stage

Almost nothing in this pipeline responds to persuasion, and almost all of it responds to evidence. The work splits cleanly by stage, and it is worth sequencing it that way rather than doing everything at once.

At retrieval, the levers are technical and linguistic: allow the search crawlers you want (checking each platform's documented user agents), keep key commercial pages indexable and current, and use the plain category words your buyers type. At source trust, the levers are external: get into the roundups and directories that already rank for your category prompts, correct stale profiles, and make sure independent descriptions of you agree with each other.

At entity resolution, the levers are structural: one naming pattern, matching contact details everywhere, LocalBusiness or Organization structured data, and sameAs links to your canonical profiles. At synthesis, the levers are editorial: state qualifying facts in the first sentence of a section, publish comparison-shaped content with tables, and give the model verifiable differentiators it can quote without risk.

None of that is a trick, and none of it is permanent. Indexes refresh, competitors publish, and assistants change retrieval behaviour without announcement, so the measurement loop matters as much as the fixes. The businesses that get recommended consistently are the ones whose evidence base is easy to find, easy to trust, and impossible to confuse with anyone else's.

See what AI says about your business

Growli measures your share of AI answers across ChatGPT, Gemini, Claude and Perplexity — and turns every gap into prioritized actions.

Get Started

FAQ

ChatGPT expands your prompt into web searches, retrieves a set of pages, and writes an answer from the ones it trusts, linking to its sources — OpenAI describes ChatGPT search as combining its model with information from the web. Because the shortlist comes from retrieved documents, third-party pages such as directories, review sites and "best of" roundups usually determine which companies are candidates, while a company's own site mainly confirms details.

You cannot buy a place in an organic recommendation. What you can influence is the evidence the assistant retrieves: whether its crawlers can access your pages, whether independent sources describe you accurately, and whether your identity is unambiguous. Some platforms have separate commercial arrangements with publishers or advertisers, but those are distinct from the synthesised recommendation itself.

Because the retrieval layer differs. Google states its AI features in Search use the same Search systems and indexing, while Perplexity retrieves through its own stack and cites sources inline, and Copilot inherits Bing's index. Different indexes hold different documents with different freshness, so the evidence set — and therefore the names — diverge, before you even account for non-deterministic generation and localisation.

It can. OpenAI documents separate bots for training and for surfacing sites in ChatGPT search, so a blanket robots.txt block written to prevent training can also remove you from the search-grounded answers. Check each platform's published user agents and decide per bot rather than blocking all AI-related crawlers by default.

Specific, verifiable, early-stated facts: the sectors you serve, deal sizes, geographies, certifications, integrations and turnaround times. Research on long-context behaviour shows models use information at the beginning and end of an input more reliably than material buried in the middle, so put the qualifying claim in the first sentence of a section and keep comparison content in clearly structured tables.

Continuously rather than occasionally, because answers change when indexes refresh, competitors publish or retrieval behaviour shifts. A single spot-check is unreliable given that generation is non-deterministic and results vary by location. Running a fixed set of buyer-intent prompts across several assistants on a regular schedule, and recording the cited sources, is the only way to see real movement — that is the loop Growli automates.

Weekly newsletter

Weekly AI search growth tips

Join business owners getting practical tips for being found on Google and recommended by AI.