A shopper no longer types "running shoes" into a search box and scrolls. They ask an assistant a full question, with a budget and a need, and they get back a handful of names. We sat down with one of those questions to show what happens next, who makes the answer, and why everyone else is invisible.
A note before we start. This is an illustrative walkthrough, not a study. The five brand names below are invented, and the pattern is a composite of what we see when we run prompts like this one during audits. There are no real brands and no measured numbers in this post. For measured results, by category and by quarter, read the Agent Index and its methodology.
What did we ask, and how?
We asked one question, written the way a shopper writes it: "What are the best running shoes under $150 for daily training?" The question has a category, a budget and a use. That matters, because each of those three parts becomes a filter the agent has to satisfy with evidence.
We used the consumer product in its default state, with no memory and no custom instructions, which is the same setup described in our Index methodology. We then asked the same question again, several times. Answers vary from one run to the next, so a single screenshot tells you very little. What you are looking for is the set of names that keeps coming back.
We also asked two follow-ups that real shoppers ask: "Which of those is best for wide feet?" and "Which one can I return if it doesn't fit?" Follow-ups are where thin product data gets exposed.
Who showed up in the answer?
Five brands made the answer, and one of them led it almost every time. In our composite they are called Kestrel Run, Northtrail, Stride Lab, Pacer & Co and Volant. None of them exist. The shape of the answer is what is real: a lead recommendation with a reason, three or four alternatives each tied to a need ("best for wide feet", "best cushioning for the money"), and nothing at all about anyone else.
That last part is the uncomfortable one. There is no page two. A brand that is not in those five names was not ranked sixth. It was not considered, as far as the shopper can tell. We call that group of names the candidate set, and getting into it is the whole game.
| Brand | In the answer? | What the agent could verify |
|---|---|---|
| Kestrel Run | Lead pick, most runs | Price, weight, drop, width options and return window, all matching across feed, page and reviews |
| Northtrail | Most runs | Strong third-party reviews; price confirmed in feed |
| Stride Lab | Most runs, "best for wide feet" | Width stated as a product attribute, not buried in a size chart image |
| Pacer & Co | Some runs | Good reviews, but the sale price on the page disagreed with the feed |
| Volant | Some runs | Known from forums; product page hard to read without scripts |
| Everyone else | Never | See the four reasons below |
Why did one brand lead the answer?
The lead brand won because every claim the agent wanted to make about it could be checked in more than one place. It was not the cheapest shoe and it was not the most famous. It was the easiest to be confident about.
Look at what the question demands. "Under $150" needs a current price. "Daily training" needs some statement of purpose: cushioning, durability, weight. "Best" needs someone other than the brand saying so. Our fictional winner, Kestrel Run, had all three lined up:
- A price the agent could trust. The product feed, the structured data on the page and the visible price all said the same number. When those disagree, agents tend to hedge or drop the product.
- Specifications as data. Weight, heel-to-toe drop, width options and intended use were written as text and as product attributes, not locked in a lifestyle image.
- Independent agreement. Buying guides and runner forums described the shoe in the same terms the brand did. The agent could quote someone else.
When we asked the follow-up about returns, the lead brand also had a return policy the agent could state in one sentence, because it was published as plain text and as return-policy markup. Two of the other four got a vaguer answer: "check the retailer's return policy". That is an agent telling the shopper it could not find out.
Why were the other brands left out?
Brands are left out for four recurring reasons, and none of them is that the shoe is bad. The agent builds its shortlist from what it can read and verify in the few seconds it has. Anything that makes a product harder to verify makes it easier to skip.
- The price could not be confirmed. The feed said one thing and the page said another, or the price only appeared after a script ran. An agent answering a question with a budget in it will not guess.
- The product was not described in the shopper's terms. The page said "engineered for the everyday athlete". The shopper asked about daily training and wide feet. If the words and attributes do not meet the question, the match is never made.
- Nobody else vouched for it. Agents lean on reviewers, buying guides and forums. A brand with no presence in the sources an agent trusts for running shoes has nothing to be quoted from.
- The store pushed the agent away. Bot walls, aggressive pop-ups and pages that render nothing without three scripts stop an agent the same way they would stop a customer with a slow phone. We cover this in why Muse reads your site like a customer.
Notice what is not on the list: ad spend, domain age, how many keywords the page ranks for. Some of the old signals still help indirectly, because a well-known brand gets written about more. But a brand with modest search rankings and clean, consistent product data can make the answer ahead of a bigger rival that is hard to read.
Does the answer change from one run to the next?
Yes, the answer changes between runs, which is why one screenshot proves nothing. In our walkthrough the top two or three names were stable and the last two slots rotated. That rotation is where most of the opportunity is. A brand that appears in some runs is already in the candidate set part of the time. Its job is to become a name the agent can be sure of every time.
This is also why we measure share of recommendation and not rank. If you ask the same forty questions five times each across five agents, as we do for the Agent Index, you get a percentage: the share of answers that recommend you. That number moves when you fix things, and it can be tracked against named competitors. A rank of "third" in one answer on one day cannot.
How can you run this test for your own category?
You can run a rough version of this test in half an hour with nothing but a notepad. It will not be rigorous, but it will tell you whether you have a problem.
- Write ten questions the way shoppers ask them. Full sentences, each with a budget, a use and a constraint. "Best trail running shoes under $130 for rocky ground" is better than "trail shoes".
- Use a clean session. Signed out where possible, no memory, no custom instructions. You want what a stranger sees, not what the assistant has learned about you.
- Ask each question at least three times, in fresh chats. Write down every brand named and whether it was recommended or only mentioned.
- Ask the follow-ups. Sizing, returns, delivery time, a comparison with a named rival. Note where the agent gets vague or gets your details wrong.
- Repeat in a second agent. Gemini, Claude and Perplexity draw on different sources. Being present in one and missing from the others is a finding in itself.
- Trace every error back to a source. A wrong price usually comes from a stale feed or conflicting markup. A missing brand usually comes from missing third-party coverage.
If you would rather not do it by hand, the free check looks at one product URL and shows you what ChatGPT says for your category. The paid Audit does this across the whole store and all five agents, with transcripts. For the fixes themselves, see our practical checklist for getting recommended by ChatGPT.
Key takeaways
- An AI shopping answer names around five brands. Everyone else is invisible, not ranked lower.
- The lead recommendation goes to the product that is easiest to verify: price, specifications and returns that match across feed, page and reviews.
- Brands are skipped for unverifiable prices, copy that does not meet the question, no independent coverage, or a store that blocks agents.
- Answers vary between runs, so measure share of recommendation across many runs, not rank in one.
- You can run a rough version of this test yourself in thirty minutes with ten well-written questions.
Questions and answers
Are the brands in this walkthrough real?
No. Kestrel Run, Northtrail, Stride Lab, Pacer & Co and Volant are invented names, and the results are an illustrative composite of patterns we see in audits. Measured results for real categories are published in the Recommended by Agents Agent Index.
How many products does ChatGPT usually recommend?
A typical shopping answer names around three to five options, often with one lead pick and the others tied to a specific need such as wide feet or best value. There is no second page, so brands outside that group are not seen at all.
Why does ChatGPT give different answers to the same question?
Language models generate answers with some randomness, and they may retrieve different sources on each run. The top names tend to be stable while the last slots rotate. That is why we measure share of recommendation across many runs instead of one result.
Does ad spend or search ranking decide who gets recommended?
Not directly. Agents favor products they can verify: a confirmed price, clear specifications, a readable return policy and independent reviews that agree. Well-known brands benefit from being written about more, but a smaller brand with clean, consistent data can appear ahead of them.
Written by the team at Recommended by Agents. Published . Spotted something out of date? Tell us and we will fix it.