How one AI citation case study is built so you can check it

The Revenue Signal publishes a company case every week. This page walks through the first one, step by step, with the actual records from the run, so a reader can see what “you can hold us to it” means in practice. This introduction offers an AI citation case study in action. The case is National Retail Solutions in the convenience-store POS category, run on 6 September 2026.

 

Step 1. The prediction is written before the data

Step 1

On 5 September the study owner wrote down the expected result and the rule for reading it. Predicted agreement between the mention ranking and the citation ranking: 0.70, range 0.55 to 0.85. Decision rule: at or below 0.70, the claim that citations reorder the picture is supported for this category; at or above 0.85, it is narrowed; between, inconclusive. Neither line was edited afterward.

Why this matters: without a prediction on paper, every result reads as confirmation. The run came in at 0.699. That is support by six ten-thousandths, and the report says so rather than rounding it into a win.

 

Step 2. The questions are brand-neutral and frozen

Step 2

Twelve questions generated against a fixed twelve-slot frame used for every company: the direct shortlist question, a feature-specific shortlist, two problem questions, three comparisons, a buy-or-not question, two how-it-works questions, a multi-store question, a US-specific one. No question names a brand, a product or a vendor. No question contains a year, so a repeat run asks exactly the same thing. One generated question contained a year and was regenerated; the rejection is logged.

 

Step 3. The brand list is built before any answer is seen

Step 3

Twenty brands with their name variants and owned domains, a written rule for names that are also common words (Square, Toast, Clover and four others), and a written rule for brands that appear unexpectedly. Three brands surfaced in the answers and were added under that rule, dated. The frozen list was not edited.

 

Step 4. Freeze, then push

Step 4

SHA-256 hashes of the question file, the brand list and the engine settings were written into the pre-registration. The run folder was committed, the commit hash was recorded in a second commit, and the repository was pushed at 21:00:25 UTC on 6 September. The first query went out at 21:03:50. The push receipt is the proof of date.

 

Step 5. The run

Step 5

Twelve questions, three assistants, three runs each: 108 planned calls. ChatGPT on gpt-5.6-luna, Claude on claude-opus-4-8, Perplexity on sonar-pro, through their APIs with web search required, so a citation was possible on every answer. 108 completed, 108 with confirmed search, 0 failed. The collection stopped on its own at 102 for a reason that was not identified; the last six were sent seven minutes later and each send time is recorded. Every raw answer is stored unedited.

 

Step 6. Scoring, with a person on the common words

Step 6

For every answer and brand, two things recorded separately: named in the text, and own page linked as a source. Third-party pages about a brand logged and never counted as its citation. Every match on a common-word brand name went to the study owner: 265 matches, each ruled against its answer.

What the scoring produced for the subject: named in 30 of 108 answers, cited from its own site in 41. Named on 7 of 12 questions, cited on 6. In 27 answers both; in 14, cited without being named; in 3, named without a citation of its own.

 

Step 7. Every cited page is opened

Step 7

320 citations, 104 pages. For each: did it load, its title, the check date, the claim the citation sat beside, and whether the page supports that claim. “Cited” and “supports the claim” are separate columns and never merged.

The first automated pass was wrong in two ways, truncating long pages and scoring unloaded pages as unsupported. It was superseded by a second pass that read every page in full, and the superseded results are kept and labelled. 54 rows were then reviewed and signed by the study owner. Of the subject’s convenience-store page, 25 citations reviewed: 13 supported, 7 partly, 3 did not, 2 not checkable. Of its EBT article, 11 of 11 supported. Rows not signed carry the automated verdict and are never called verified.

 

Step 8. Guards before statistics

Step 8

Sparsity guard: at least four brands must have owned citations in three or more answers, or the verdict is “insufficient citation data.” Observed: eleven. Passed. Power: at least ten ranked brands for the decision rule. Observed: fourteen. Passed.

 

Step 9. The statistics, with a hand check

Step 9

Spearman’s rho between the two rankings, pooled across the three assistants: 0.699. Prompt-level bootstrap interval, 2,000 draws: 0.475 to 0.873. Per assistant: ChatGPT 0.873, Claude 0.643, Perplexity 0.323. The pooled figure was recomputed by hand in the report’s appendix and agrees with the library value to four decimal places.

 

Step 10. The report, then the write-up

Step 10

The report is written in a fixed order, guard result first. The published write-up is built from the report and nothing else. If a claim is not in the report, it is not in the write-up.

 

What this run found

What this run found

The subject is the most-cited brand in the category and absent from five of the twelve questions. In 50 of 108 answers a competitor’s own page was cited and the subject was neither named nor cited. The pages that took those citations share one build: the question as the title, a figure or definition in the first screen, a table, a numbered procedure. The subject’s own EBT article has that build and won 11 of 11. Its convenience-store page does not, and is cited for specifics it does not carry.

 

What it will not claim

What it will not claim

One category, one day, counts not rates. No ranking, citation count or revenue figure is guaranteed. Nothing here observes buyers or sales. The pattern across nine pages is consistent with published research and is not a controlled test of any single page.

 

What subscribers get

-what-subscribers-get

The public case ends at the diagnosis. The Revenue Signal carries the prescription: for this run, five recommendations, each tied to a count above, on what to write, what to put on the page that is being cited for things it does not say, how to get the brand into the sentences that get lifted, why one assistant needs its own plan, and how to measure the change. That section exists only in the newsletter, and the categories come from readers. Reply with yours and I will tell you where you stand.

Subscribe here:

To have the method run on your own category, at either depth, book a fit conversation with John Bush: https://meetings.hubspot.com/john2956

Evidence package: pre-registration with dated amendments, frozen questions, brand lists, engine settings, freeze commit df02447, derived answer records, page-check table and report, exported from the private repository: https://drive.google.com/drive/folders/1UP0g6SyUt2IVpn1vl3R-FcxHn5KSMX0B?usp=sharing . The full text of each engine answer is withheld pending a check of each provider’s terms on republishing outputs. Method: RE-METH-AIV-003.

Discover more from Revenue Experts AI

Subscribe now to keep reading and get access to the full archive.

Continue reading