Being named by an AI is not the same as being recommended by one.
Elizabeta Kuzevska, Revenue Experts AI
Your buyers have not stopped comparing vendors. They have changed where they do it. Before anyone reaches your site, some share of your shortlist is now being drawn up by an AI system answering a question like “best option for a business my size.” So teams have started checking their AI visibility, and almost all of that checking measures one thing: does our brand get mentioned.
The answer usually comes back reassuring. It is also usually the wrong measurement.
Being named by an AI is not the same as being recommended by one. A model can produce your company’s name from memory, then cite a competitor’s page, or a review site, or a clinic, as the source for what to actually buy. Count mentions and that gap is invisible to you. It is not invisible to the buyer, who is reading the recommendation, not the mention.
I publish a teardown every week that takes this apart on a real company. One buying category. One company used as the case. The specific pages the engines cited instead of theirs. It runs in The Revenue Signal, and this piece is the method behind it, written so you can hold us to it. A measurement you cannot examine is a marketing claim.
A teardown starts with the questions a buyer types when they are choosing, none of which name a brand. It runs those questions through the answer engines your buyers use, ChatGPT, Claude and Perplexity, with web search on, so that a citation was possible on every answer. Then it reads the results the way an engineer reads an engine you can test but cannot open.
Most of the market stops at a count of appearances. The method behind these teardowns does six things that a mention counter does not.
It counts four events, not one. Four different things happen inside an answer, and they carry different commercial weight. Your brand named in prose with no source. Your own page cited as a source. Someone else’s page about you cited. And a citation whose page actually supports the claim it was attached to. Those four never get collapsed into a single “visibility” number. A company can score well on the first and nothing on the fourth, and only the fourth is evidence.
It asks more than once. Engines answer the same question differently from one session to the next. So every question runs several times, each a fresh session, each stored separately rather than averaged away at the moment of collection. One appearance is a sighting. The share of runs a brand actually held is a rate. Those are different claims, and the method keeps them apart.
It opens every cited page and checks it. This is the part almost nobody does. A citation tracker records that an engine linked a page. It rarely opens that page to ask whether the page supports the specific sentence the engine attached to it. An engine can cite something adjacent, outdated, or flatly contradicted by the page itself, and a counter files all three as wins. Every cited page in scope gets opened, and “cited” and “supports the claim” are recorded as two separate results.
A person signs off what matters. The low scores, the citations pointing at a client’s own domain, the priority findings, those go to a named reviewer before they become findings. The model is not the final judge. On our own audit, that review caught the automated pass understating page support, in the direction that would have flattered nobody. The method caught it because a human had to sign each row.
Every finding is labelled by the kind of evidence it is. A measured fact from the run, a pattern seen in the sample, an untested theory, and a result from a controlled test are four different things. They get named as four different things, so you always know which one you are reading and what its limits are.
The prediction is written before the data, and the design is frozen before the first question. The expected result goes on paper, dated, with the rule for what each outcome will mean, and it is not edited afterwards. The questions, the brand lists and the settings are committed and time-stamped to a public receipt before a single query is sent. Without a prediction on paper, every result reads as confirmation. A prediction that turns out wrong is a result, not an embarrassment. It is the only structure that lets a run disagree with the person who ran it.
That is the whole discipline in one sentence: every figure a teardown reports traces back to a design that was committed before the data existed, and to raw answers stored unedited. You do not have to take my word for the number. You can follow it back to where it came from.
I would rather set the limits myself than have you find them.
A weekly teardown is a first look, not a verdict. It runs a category once. It can show you a gap. It cannot tell you, on its own, how stable that gap is over time, because that takes repeated cycles and a control set to separate the engine’s own drift from anything you changed. Two readings a month apart are not a trend. Comparing two sightings and calling the difference a result is the most common error in this category, and it is one we do not make.
The deeper work is a paid study on your own category: the same method run wider and on a schedule, every cited page checked with human sign-off, and where you need cause rather than correlation, a controlled test with matched pages, a baseline, and a difference-in-differences read on what actually moved.
Here is the honest position on that. We have run the full method on our own domain: multiple engines, repeated runs, every citation scored against the sentence it was attached to, and a report generated from it. That scoring is an automated first pass. The human review of the flagged citations is queued and not yet worked. No paid client engagement has been through the full method. What I will not promise is a citation count, a ranking position, or a revenue figure. No platform publishes its formula. What is guaranteed is the work and the evidence behind it.
Because the ground moves weekly. Brand presence in AI answers is comparatively durable. The evidence layer underneath it, the pages the engines actually pull from to justify a recommendation, gets renegotiated far faster. A company can hold its spot in the consideration set while competitors quietly take over the sources beneath the recommendation, and a share-of-voice report will show that as stable performance.
There is also a reason it belongs to more of you than the coffee brand in any given week. 94% of B2B buying groups rank their shortlist by preference before contacting any vendor (6sense, 2025). If that ranking is increasingly assembled by a system reading pages you may not have written for it, the question of which pages it reads stops being a marketing detail. It is a revenue question.
Each teardown is one category on one date, stated as exactly that. Read enough of them and the pattern underneath is the useful part: the mention-versus-citation gap is diagnostic, your competition for a citation is usually not your competitor, and the fix is almost never more brand content.
You can run the first check yourself, and it is free. Take one question your buyers actually ask, with no brand name in it, and put it to ChatGPT, Claude and Perplexity with search on. Note whether they name you. Then note whether they cite you. Treat those as two separate answers. If you are named and not cited, more brand content will not close it. Then read the pages that got cited instead of yours, and ask what they answered that your page argued.
That diagnostic is the thing the weekly teardown does out loud, on someone else’s company, so you can see the method before you ever pay for it.
Every week in The Revenue Signal I do this on a real company, in a category that could be yours next. One buying category. One company used as the case. The specific pages ChatGPT, Claude and Perplexity cited instead of theirs, with the counts and the date.
Subscribers get the part I keep out of the public version: the recommendations. What the cited pages did that the company’s page did not, and what to change so a retrieval system can actually use yours. That is the piece you can act on before the next issue lands.
Two more things come with it. A self-check, so you can put your own mention-versus-citation gap on paper in one afternoon. And a reply line. Reply with your category and I will tell you where you stand. I read every reply, and reader categories become future teardowns.
Here is why it is worth doing now rather than later. The name you hold in AI answers is fairly durable. The evidence underneath it, the pages an engine leans on to justify a recommendation, gets renegotiated week to week. A competitor can take over the sources beneath your recommendation while your mention count stays flat. You will not see that in a share-of-voice report. You will see it in a teardown.
What it will never do is promise a ranking, a citation count, or a revenue figure. You get counts with dates, every cited page opened and checked, and the limits stated every time.
And if you would rather have this run properly on your own category than watch it on someone else’s, that is a conversation. Book a call with John Bush to scope it: meetings.hubspot.com/john2956
Subscribe now to keep reading and get access to the full archive.