How the Index is made.
Forty questions, five agents, two markets, one week. Here is exactly what we ask, how we ask it and how we turn the answers into a ranking.
“What are the best running shoes under $150 for wide feet”
What the agent said
- 1Brand ARecommended1 point
- 2Brand BRecommended1 point
- 3Brand CPassing mention0 points
- Coded by two people independently
- Transcript and screenshot kept
- Questions per category
- 40
- Agents equal weight
- 5
- Runs per question, per agent
- 5
- Answers per category, per market
- 1,000
What we measure
Share of recommendation: the percentage of answers in which an agent recommends a brand when a shopper asks for the best product in a category.
We also record candidate-set presence, which is whether the brand is mentioned at all, and position, which is where it appears in the answer.
- Share of recommendation
- The ranking is built on this.
- Candidate-set presence
- Whether the brand is mentioned at all.
- Position
- Where it appears in the answer.
One edition, start to finish.
Every category goes through the same five steps, inside the same seven-day window.
-
Ask
Forty questions per category. They are written as shoppers write them: full sentences, with a budget, a use and a constraint.
A quarter of the questions change each edition so that nobody can tune for the test. The rest stay fixed so that results compare across quarters. We publish the full list with each edition.
“What are the best running shoes under $150 for wide feet”
A typical question.
- A quarter change each edition
- The rest stay fixed
-
Run
ChatGPT, Gemini, Meta's Muse, Claude and Perplexity. We use the consumer product in its default state, signed out where that is possible, with no memory and no custom instructions.
Each question is asked from a UK and a US location, and the two markets are ranked separately.
Every question is asked five times per agent, because the answers vary from one run to the next. That is one thousand answers per category per market.
All runs happen inside the same seven-day window. We keep the full transcript and a screenshot of every answer.
Five agents
- ChatGPT
- Gemini
- Meta's Muse
- Claude
- Perplexity
Two markets
- UK
- US
Five runs per question
One seven-day window
-
Read
Two people code every answer independently. Where they disagree, a third decides.
For each brand in an answer we record whether the agent recommends it, whether it is mentioned at all, and where it appears.
- First coderRecommended
- Second coderPassing mention
- Third decidesPassing mention
Example disagreement, for illustration only. -
Score
A brand scores one point for each answer in which the agent recommends it, and nothing for a passing mention.
Share is points divided by total answers. All five agents carry equal weight.
Sharepoints divided by total answers
Example shares, for illustration only. -
Publish
The Index is published quarterly, with the UK and the US ranked separately and the full list of questions alongside.
Nobody sees the results before they are published.
Published each quarter
- The UK ranking
- The US ranking
- The full list of questions
- Clients marked as clients
What this does not tell you
It does not measure sales, traffic or how good a product is. It measures what the agents say.
- SalesNot measured
- TrafficNot measured
- How good a product isNot measured
- What the agents sayMeasured
Agents change without notice, so a ranking is a record of one week, and the trend across quarters matters more than any single position.
Independence
Our clients are ranked the same way as everyone else.
- Nobody can pay for a position.
- Nobody sees the results before they are published.
- Clients that appear in a category are marked as clients.
- We publish the full list of questions with each edition.
One partner makes you findable.
Start with the free check, or book a thirty-minute readiness call.