How the Index is made.

Forty questions, five agents, two markets, one week. Here is exactly what we ask, how we ask it and how we turn the answers into a ranking.

One answer, coded Running shoes

“What are the best running shoes under $150 for wide feet”

One of five agentsUSRun 3 of 5

What the agent said

  1. 1Brand ARecommended1 point
  2. 2Brand BRecommended1 point
  3. 3Brand CPassing mention0 points
  • Coded by two people independently
  • Transcript and screenshot kept
Example answer, for illustration only.
Questions per category
40
Agents equal weight
5
Runs per question, per agent
5
Answers per category, per market
1,000
Markets
UK, US ranked separately
Window
Seven days all runs inside the same week
Published
Quarterly with the full list of questions

What we measure

Share of recommendation: the percentage of answers in which an agent recommends a brand when a shopper asks for the best product in a category.

We also record candidate-set presence, which is whether the brand is mentioned at all, and position, which is where it appears in the answer.

Share of recommendation
The ranking is built on this.
Candidate-set presence
Whether the brand is mentioned at all.
Position
Where it appears in the answer.

One edition, start to finish.

Every category goes through the same five steps, inside the same seven-day window.

  1. Ask

    Forty questions per category. They are written as shoppers write them: full sentences, with a budget, a use and a constraint.

    A quarter of the questions change each edition so that nobody can tune for the test. The rest stay fixed so that results compare across quarters. We publish the full list with each edition.

    “What are the best running shoes under $150 for wide feet”

    A typical question.

    • A quarter change each edition
    • The rest stay fixed
  2. Run

    ChatGPT, Gemini, Meta's Muse, Claude and Perplexity. We use the consumer product in its default state, signed out where that is possible, with no memory and no custom instructions.

    Each question is asked from a UK and a US location, and the two markets are ranked separately.

    Every question is asked five times per agent, because the answers vary from one run to the next. That is one thousand answers per category per market.

    All runs happen inside the same seven-day window. We keep the full transcript and a screenshot of every answer.

    Five agents

    • ChatGPT
    • Gemini
    • Meta's Muse
    • Claude
    • Perplexity

    Two markets

    • UK
    • US

    Five runs per question

    One seven-day window

  3. Read

    Two people code every answer independently. Where they disagree, a third decides.

    For each brand in an answer we record whether the agent recommends it, whether it is mentioned at all, and where it appears.

    • First coderRecommended
    • Second coderPassing mention
    • Third decidesPassing mention
    Example disagreement, for illustration only.
  4. Score

    A brand scores one point for each answer in which the agent recommends it, and nothing for a passing mention.

    Share is points divided by total answers. All five agents carry equal weight.

    Sharepoints divided by total answers

    • Brand A41%
    • Brand B27%
    • Brand C12%
    Example shares, for illustration only.
  5. Publish

    The Index is published quarterly, with the UK and the US ranked separately and the full list of questions alongside.

    Nobody sees the results before they are published.

    See the latest Index

    Published each quarter

    • The UK ranking
    • The US ranking
    • The full list of questions
    • Clients marked as clients

What this does not tell you

It does not measure sales, traffic or how good a product is. It measures what the agents say.

  • SalesNot measured
  • TrafficNot measured
  • How good a product isNot measured
  • What the agents sayMeasured

Agents change without notice, so a ranking is a record of one week, and the trend across quarters matters more than any single position.

Independence

Our clients are ranked the same way as everyone else.

  • Nobody can pay for a position.
  • Nobody sees the results before they are published.
  • Clients that appear in a category are marked as clients.
  • We publish the full list of questions with each edition.

One partner makes you findable.

Start with the free check, or book a thirty-minute readiness call.