How Perplexity, ChatGPT And Gemini Pick Their Sources
How to Test Rather Than Trust Everything above is a starting hypothesis. Run twenty prompts in your own category across all three, from signed out sessions, recording the mode and the date, and count the cited domains for each.
If the baseline exists, access problems were found and fixed, listings were corrected with names attached, and the source list has begun to move, the engagement is on track even if mention rate has not shifted. If none of those happened, the next ninety days will not be different from the first. ai overviews optimization
Comparison Is the Native Format Shopping questions are comparison questions. Somebody asking what to buy wants options weighed against each other, so the sources that get used are the ones that have already weighed them.
Weeks Three and Four: The Access Findings A technical report covering crawler permissions, what the relevant agents actually receive from your server, whether bot management is interfering, and what survives on your key pages with JavaScript disabled.
That means the useful ask is not simply for a rating. Prompting customers to say what they used the product for and what situation it suited produces review text that can actually answer a question, which is what gets quoted.
Assume the pitch is good. Everyone's pitch is good, and the vocabulary in this field is easy enough that a competent salesperson can hold a convincing conversation without anyone behind them who can do the work.
How to Run the Ninety Day Review Ask three questions. Can you show me the prompt set is unchanged. Can you show me the raw answers. What specifically did you do, and which of the changes do you believe caused which movement.
Long sections on activity that produced nothing, described in the language of effort rather than outcome. And the most reliable indicator, a report you cannot disagree with, because it contains no specific claim to test.
Most engagements are judged too late, on a final outcome that arrives after the point where anything could have been corrected. The first quarter has its own deliverables, and knowing what they are lets you tell early whether you have hired the right people.
You are unlikely to read all of it, and its presence changes the incentives entirely. An agency that knows the raw evidence ships with the report writes a different summary than one that knows it will not be checked.
Finally, pay attention to how they talk about their existing clients. Somebody who describes a client's category accurately, names the specific constraint that made the work difficult, and mentions something that did not work has actually done the job. Somebody who describes every engagement as a success in identical language has either been unusually lucky or is describing a template.
This frequently produces the first result of the engagement, because access failures are total and fixing them can change answers within days. It should also be short. A fifty page technical audit at this stage is usually padding drawn from a generic template.
Ask who specifically will do the work, and ask to meet them. Capability in this field is thinly distributed and frequently sits with one or two people inside an agency of any size. A pitch delivered by a strategist who then hands the account to a junior is common enough to be worth checking for directly, and the question is easy to ask without giving offence.
Product recommendations are a harder case than service recommendations, because the answer has to be specific enough to act on. A model naming a product is committing to a name, usually a price band and often a comparison, and it needs sources confident enough to support that.
The Shared Architecture All three now commonly retrieve live sources rather than answering purely from training. Your question becomes one or more searches, a set of pages is fetched and read, and the answer is composed from what was read.
Refusals matter. A report that only contains successes is either describing a suspiciously easy month or omitting the parts that did not work, and the omitted parts are usually where the useful information is.
Observed behaviour leans toward breadth, pulling from a wider set of sources per answer than the others, and it cites forums, documentation and niche trade sources readily. It also appears comparatively responsive to freshness.
Ask how they will handle being wrong. Every engagement in this field produces at least one confident recommendation that does not work, because the systems change and the published research is thin. What matters is whether that gets reported or quietly dropped from the next deck, and asking the question directly at the outset makes it considerably more likely to be reported.
Two implications follow regardless of which system you are studying. Being findable by the underlying search step is necessary, and being worth quoting once fetched is what decides whether you are used. Almost everything actionable sits in those two requirements.