How Algotally measures AI recommendations
What we ask, how often, what we count, and how sure each number is.
Why this page exists
Ask an AI assistant the same question twice and you will often get a different list.
SparkToro's 2026 study
found the same list of brands came back in under 1% of repeats. The set of brands AI draws from
was much steadier, though. So a single answer, or a "rank #3", says very little. How often a business
is named across many answers says a lot more. Every number Algotally reports is that second kind,
and each one comes with how many answers it rests on.
What we ask
- For each market (an industry in a city), we ask a small fixed set of buyer-intent questions,
e.g. "I need a plumber in Houston, TX. Which local companies do you recommend?"
- We ask the same questions again and again rather than many different ones. That
measures how consistently a business is named, not whether it turned up once.
- Engines: Perplexity, OpenAI (ChatGPT's models) and Anthropic Claude, each with live web search
switched on. Each engine is asked up to 8 times per market per week, and up to
24 in markets an Agency customer follows, which tightens the ranges below for everyone there.
- We also record a separate "from memory" check with web search off. It is never mixed into the
numbers above, because it measures something different.
What we count
- Mention rate: the share of answers that named the business.
- Share of voice: the business's slice of every recommendation AI made in the market.
If answers named businesses 200 times in all and 20 of those were this business, its share of voice is 10%.
A business is counted once per answer, however many times the answer repeats its name.
- Refusals and empty answers are left out of the count. An engine declining to answer
is a missing observation, not evidence against a business.
- Cited sources: the web pages the engine cited before answering. A source is only
ranked once it has been cited in at least 2 different answers, and rankings use
the last 90 days only.
- Model changes are flagged. AI providers swap the model behind a product without
notice. We record which model actually answered, and mark any week where that changed. A jump
that week may be the model, not the business.
How sure each number is
- Every mention rate comes with a 95% range (a Wilson score interval). For example,
named in 6 of 24 answers is 25%, with a range of 12%–45%.
Named in 0 of 24 still leaves a range of 0–14%: "not seen" is not "never".
- Under 5 answers we show no range at all. That is too few to say anything.
- Week-to-week moves are tested. A change is called real only if it clears a 95%
two-proportion test, with at least 10 answers on each side. Otherwise it is marked
within noise.
- Honest caveat: these tests treat answers as independent. Repeats of the same question are
somewhat correlated, so true uncertainty is a little wider than shown. Read borderline results
with that in mind.
Before and after
When a subscriber logs an action (e.g. "got listed on a directory"), we compare the mention rate in
up to 4 weeks before with up to 8 weeks after, skipping the week of the change itself.
The result reads "too early" until 2 weeks of data exist afterward. It gets the same noise test as
above, and is flagged if the models changed between the two windows. This shows what followed an action,
not proof of what caused it.
What we don't do
- We don't give any business or source a quality score. A source ranks high here because AI
cites it, not because we endorse it.
- We don't track geographic grids or personalised answers. Questions name the city directly.
- We don't show a single answer as if it were typical.