Your AI visibility number was never checked against the page | Revenue Experts AI

A question worth asking your provider today

Has anyone actually opened the pages your AI visibility score is built on?

Almost certainly not. Every tool in this category produces its number the same way. It sends a prompt, reads the answer that comes back, and counts brand names and domains inside that answer. It never fetches the page the answer pointed at.

The score describes what the engine said. Not what the source says.

That gap is where the reason you are not cited lives.

Before the argument startsThe Revenue Signal is the Revenue Experts AI weekly read for B2B revenue leaders. Every Thursday, one shift in how buyers research vendors, the data behind it, and one move you can make before the next issue lands. No hot takes, no vendor pitches dressed up as insight. If how AI visibility actually gets measured matters to you, subscribe here.

On 3 August 2026 the IAB published a thirty-six page measurement framework called Measuring Visibility in the AI Era, built by a working group drawn from Walmart, Acxiom, WPP Media, Microsoft Clarity, EMARKETER, PMG and Tinuiti.IAB pg.8

Buried in the metrics section is a requirement that quietly breaks the standard way of measuring.

The framework defines two separate failures a provider must detect and report. A hallucinated mention is one the engine fabricated, with no basis in any underlying source. A factual inaccuracy is different: the engine names you correctly, but attaches materially wrong information, because it is faithfully reflecting a source and the source is wrong.IAB pg.16

Providers must disclose their detection methodology for both, report them per platform rather than blended, and surface flagged mentions to clients rather than quietly dropping them. The IAB tells buyers to treat a missing detection methodology as a material gap in any provider's quality claim.IAB pg.22

Now think about how you would detect a factual inaccuracy.

You cannot find it by scanning the answer text, because in the answer your name is spelled correctly and the sentence reads fine. That is precisely what makes it a factual inaccuracy rather than a hallucination. The only way to find it is to go to the page the engine cited, read what is on it, and compare that against what the engine claimed it said.

Somebody has to fetch the source. A system that only reads answers cannot do it, no matter how many prompts it runs or how much it costs.

That is the gap this page is about. It is not that your dashboard is inaccurate. It is that your dashboard is describing a different object than the one you need described.

I will show you exactly what checking the source produces, and what it finds that counting names never can. First, the thing that goes wrong one step earlier.

Argument one

Before anyone checks the page, most tools have already guessed what the citation was for

When an AI engine cites a source, the citation is attached to something. A sentence. A claim. A specific assertion inside the answer.

Knowing which one matters enormously. "Cited alongside a paragraph about pricing" and "cited as the source for the pricing figure" are different events with different fixes.

Here is how most tools establish that link. They take the answer text and the list of cited URLs, hand both to a second language model, and ask it to work out which claim each citation was probably supporting.

That is a guess. A reasonable one, often. But the precision figure a tool reports afterwards is not measuring the engine's citation. It is measuring how well a second model reconstructed a link it never saw.

The engines already tell you, if you read the right field

Claude, ChatGPT and Perplexity all emit the binding themselves. Claude returns a cited_text field. ChatGPT returns the answer span the citation attaches to. Perplexity returns the [n] marker in position. The information is in the API response. It does not need reconstructing.

Revenue Experts AI built this pipeline the wrong way first, with a second model re-deriving claims. Two numbers from that build are worth stating plainly, because they are the reason this page exists.

With re-derivation, only 12% of citations ended up with a claim attached at all, according to the Revenue Experts AI build record for the June 2026 validation run. The precision figure the system produced in that same June 2026 run was 14%, according to that same Revenue Experts AI build record, which we later established was not a real measurement of anything. It was an artefact of the reconstruction failing.

Switching to the engines' native binding took claim population from 12% to 100% on the June 2026 re-run, according to the Revenue Experts AI build record, and turned precision into a figure that describes the actual citation.

The Revenue Experts AI build record calls re-derivation the single load-bearing risk in the entire system. It was, and Revenue Experts AI is the firm that got it wrong first.

Then there is the sampling problem, which the IAB has now written down

Even a correctly bound citation is one draw from a distribution. The framework is direct about it: single-response measurement is not measurement, a brand's visibility on a given query is a distribution rather than a value, and any metric derived from one response per query reflects a sample of one. It adds that this is how large language models generate text and it cannot be engineered away.IAB pg.32

Independent work put numbers on it before the framework existed. In a January 2026 SparkToro study, six hundred volunteers ran twelve prompts through ChatGPT, Claude and Google's AI a combined 2,961 times. There is less than a one in a hundred chance of the same brand list appearing in any two responses, and closer to one in a thousand for the same ordering.SparkToro

Roughly half of all cited domains change every month, per Similarweb data, and only eleven per cent of citations overlap between ChatGPT and Perplexity, which leaves eighty-nine per cent of citation opportunities specific to a single platform.Similarweb

Across 126 million United States AI search prompts collected between January and April 2026, per Semrush, only thirty-six of more than twelve hundred tracked brands held top-hundred visibility on every platform in every month of the window.Semrush

And your number can move with nothing on your side changing. The framework states that share of voice can shift five points overnight because a model was retrained, with no change to the brand and no change to its competitive position.IAB pg.33

Which is why the four obvious fixes are not fixes

A better tool. Non-determinism sits in the models, not the tools. A second dashboard is a second draw from the same distribution, produced the same way, from the answer text.

A more expensive tool. Price buys prompt volume, engine coverage and interface. It does not buy a fetched source or a native binding.

More optimisation. A survey of forty-five studies published on arXiv in July 2026, per the authors, concluded that experiments provide strong evidence a document already inside an engine's context can influence how it is cited, but establish far less often that a page will be retrieved at all, and almost never that it produces a durable effect on conversions. Across 171,003 documents and 2,700 queries, body-only optimisation reduced average top-twenty presence by about nine per cent, top-ten presence after reranking by sixteen per cent, and final citation by six per cent.arXiv 2607.14035

Asking the market who to trust. In a July 2026 audit published by an independent practice lab, a classification of all ninety-eight results Google ranked for "answer engine optimization services" on 10 July 2026 found sixty were a firm's own service page and twenty-two were best-agency roundups. Exactly one was a buyer asking whether any of it was worth paying for. Zero were independent analyses.SERP audit

What this leaves you holding

In the 2026 State of AEO survey, Minuttia reported that eighty-two per cent of organisations have implemented answer engine optimisation to some degree. Around forty per cent have no dedicated budget for it. Nine per cent have a dedicated specialist, and fifty-two per cent hand it to the search team on top of an existing job.Minuttia

In Fractl’s Q2 2026 survey of 150 marketers, sixty-one per cent report some confidence in their strategy, twelve per cent are very confident with measurable results, and the remainder, seventy-six per cent in total, are running these tactics with limited or no proven attribution.Fractl

Three quarters of the market is doing the work and cannot prove it did anything. That is not a content problem. That is what happens when the instrument never looks at the thing it claims to measure.

Argument two

What a pipeline looks like when it reads the binding and fetches the source

The Revenue Experts AI citation pipeline runs sixteen stages. Fourteen of them run on every engagement. Two are human by design, and they are human because a machine cannot make those two calls defensibly.

01

Generate prompts

Buyer questions across ten prompt types, each tagged with funnel stage, role, industry, geography and topic. On the Revenue Experts AI reference run of June 2026, fifty prompts, seventy-two per cent of them non-branded.

02

Prompt QA gateHuman

A person screens every prompt before a single paid query runs. Is this a real buyer question. Is it too broad. Is it a duplicate. Is it leading toward the client. Is it a genuine buying stage. Hard stop until approved. On the Revenue Experts AI reference run this caught forum venting mistaken for buyer voice, an embedded price, an out-of-scope subdomain, and an unverified claim about a competitor.

03

Run the queries

Each prompt to each engine with live web search on, three times, every run stored separately so variability can be measured rather than averaged away. Reference run: 450 of 450 completed, resume-safe through three interruptions.

04

Extract citations and native claim binding

Claim taken from each platform's own response structure. Claude's cited_text, ChatGPT's answer span, Perplexity's marker position. No second model re-deriving what the citation was for. URLs redirect-resolved, dated and classified.

05

Fetch the source and score supportHuman

Each cited page is retrieved and scored on whether it supports the bound claim, on the four-point scale set out in the Citation Audit Method. Produces Citation Precision and Support Rate, always reported per platform. A person reviews every score of zero or one, and every claim that will reach the client. Reference run: 3,696 citations scored.

06

Crawl sources

Every cited page and every client page fetched for comparison. Raw HTML, clean markdown and schema types. Fetched directly rather than through an HTML parser, because parsers produce schema false-negatives. Around eighty-four per cent read clean on the reference run.

07

Page features

What cited pages contain that non-cited pages do not. Direct answer, tables, data, named author, schema, compared as cohorts.

08

Embed corpus

Paragraph-aware chunking into a vector store with a client-page flag on every chunk, so gaps can be found semantically rather than by keyword.

09

Metrics and patterns

Ten core metrics computed over the full run. Every finding graded by what kind of evidence supports it.

10

Client gap analysis

For every prompt where a competitor was cited and you were not, the specific fixable difference. Where there is no competitor either, it is flagged as open territory rather than as a failure. Where the gap turns out to be knowledge the business holds but has never published, that is a data readiness problem rather than a content problem.

11

Generate briefs

Gaps clustered into page families, then a schema kit, an llms.txt, prescriptive page briefs and a sequenced roadmap.

12

Build deliverable

One self-contained interactive report. Per-platform precision rendered separately, never pooled. Validation table shows a stratified sample across all four support scores, including the bad ones.

13

to 16 — Controlled experimentBuilt, not run

These stages exist and are deliberately not run in a standard engagement. They require editing a live page, holding a matched control, and waiting weeks for re-crawl. Running them would earn genuine controlled-test findings. Not running them means the study is observational, and we label it that way rather than implying causation we did not test.

Shaded stages are the two human gates. They are the reason the output is defensible, and they are the reason this is not a subscription.

Three platforms, and one we will not touch

Revenue Experts AI runs Claude, ChatGPT and Perplexity. We do not run Gemini, and the reason is legal rather than technical. Google's grounding terms prohibit extracting, storing and analysing grounded citation URLs, identically on the consumer API and on Vertex. The code to query Gemini exists in our pipeline and is switched off by configuration.

For Perplexity we checked the API terms and acceptable use policy directly before including it, and use citation URLs in paid deliverables under the two conditions those terms set: cite the underlying source websites, and keep the AI-provenance disclosure.

You are entitled to know which engines a provider covers and why. That is one of the eleven disclosures the IAB now names as required.IAB pg.27

What a completed run produces

Every figure in the table below comes from a single Revenue Experts AI reference run completed in June 2026 against revenueexperts.ai, our own domain, used as the test subject while the pipeline was being validated. The full self-audit, including the pattern it exposed, is published separately. They demonstrate what the output looks like. They are not a client result and Revenue Experts AI does not present them as one.

Revenue Experts AI reference run, revenueexperts.ai, June 2026. 50 prompts, 3 engines, 3 repetitions, 450 runs.
MetricValueReading
Total citations4,322across 1,410 unique URLs
Citation rate7.6%cited in 34 of 450 runs
Citation stability92%of cells where the brand appeared, it appeared in all three repetitions
Precision, Claude78.4%tight sentence-level binding
Precision, ChatGPT62.1%answer-span binding
Precision, Perplexity33.4%low precision with a normal support rate, read together
Third-party dependency50.1%half of visibility sits on pages the brand does not control
First-citation share64.7%when cited, usually near the front of the answer
Every value above has its source in one Revenue Experts AI reference run against revenueexperts.ai, June 2026, and is not a client result.

Why we publish the 33.4%

Revenue Experts AI publishes it because pooling the three engines into one number would hide it, and hiding it would be wrong. Perplexity's precision is less than half of Claude's, so a pooled figure would land in the sixties and nobody would ask a question.

The three engines bind citations differently, as the Citation Audit Method write-up explains. Perplexity hangs several sources off one broad sentence, so each source only partly supports it, which produces low precision alongside a perfectly normal support rate. That is a property of how Perplexity cites, not a verdict on the brand.

Reporting per platform is the only honest option, and it happens to be exactly what the IAB requires at decision-grade: per-platform results reported separately, with the variance across platforms shown.IAB pg.24

If a provider hands you one blended visibility score, ask what the spread across engines was. The blend is where the interesting number goes to hide.

And why 92% is the figure that matters

Citation stability is what you get when you run every prompt three times instead of once and keep the runs separate. It answers a question a single-run score cannot even pose: when this brand appeared, did it appear reliably, or did it appear once out of three?

Ninety-two per cent means the appearances were mostly solid rather than lucky. That is the difference between a position you can build on and a coin flip you mistook for a trend.

It is also the only defensible basis for judging next quarter's number against this quarter's, which is the thing the IAB says trend interpretation is guesswork without.IAB pg.32

The meeting this changes

Picture the next budget conversation

Your CFO asks the question they always ask. Is this working.

Right now that question opens a trapdoor. You have a score. It moved, or it did not, and you can explain the methodology up to the edge of what your provider has told you. Past that edge is a shrug with a screenshot attached.

Now picture answering it with the run behind you.

You say which engines were queried, on which dates, at which model versions, how many prompts, run how many times each. You say the brand appeared in a stated share of runs and that ninety-something per cent of those appearances held across all three repetitions, so this is a position rather than noise. You say precision differed sharply by engine and here is why, and here is the per-engine breakdown rather than one blend. You say half your visibility sits on pages you do not own as of the run date, and here are those pages, by URL.

And when someone asks which pages to fix, you do not describe a gap. You hand over a list, in priority order, with what to change on each and the reason it is on the list.

Nobody in that room has to trust you. They can open the pages.

My name is Elizabeta Kuzevska, and I am co-founder of Revenue Experts AI.

Revenue Experts AI built the pipeline described above. My co-founder is John Bush. We built it the wrong way first, discovered our precision number was measuring our own reconstruction rather than the engines' citations, and rebuilt stage four around the native binding each platform already provides.

I am not going to tell you our numbers are exact. The IAB is clear that nobody's are. What I will tell you is which engines we query, how many times, what we found when we opened the pages, and where our own figures are weakest. The 33.4% Perplexity precision figure, sourced to our own June 2026 reference run, is on this page for that reason.

What we do not claim

Revenue Experts AI numbers are not exact. Non-determinism applies to our measurement exactly as it applies to everyone else's. What we claim is that we report per platform, state the repetition count, and measure stability instead of hiding it.

This is an observational study, not a controlled test. Stages 13 to 16 would run a controlled experiment with a matched control page. We do not run them in a standard engagement, so nothing in your report will claim that a change we recommend caused a citation. It will tell you what was true on the dates of the run, and what the evidence for it is.

We do not measure implied citations. Where content appears drawn from a source without explicit attribution, the IAB excludes it because it cannot be measured consistently, and so do we.IAB pg.17 Your real influence is larger than what anyone can observe, including us.

This is not a live dashboard. The audit is a point-in-time diagnosis. It establishes a baseline and tells you what to change. It is not a trend instrument, and one audit should not be used to judge whether a programme is working over time.

Proven means proven on the data tested. These engines change. Our internal standard for the word is demonstrated to work correctly on the run we did, not correct forever. The three-repetition design exists precisely because that is true.

When you should not buy this

Revenue Experts AI is the wrong purchase for some readers. If your bottleneck is on-page work you can do inside your own content management system, do it yourself. Answer blocks, question-phrased headings, a robots file, structured data. Your team can learn that in a week, and a service billing a retainer for it is selling you rebranded search optimisation.

If your bottleneck is that nobody has ever opened the pages your category is actually being cited from, that is what we are for.

The honest test is one question. Do you know which specific URLs are producing the citations you are not getting?

If yes, you do not need us. If no, no dashboard on the market is going to tell you, because that answer is not in the answer text.

The offer

The AI Citation Audit

The sixteen-stage pipeline above, run against your category, across Claude, ChatGPT and Perplexity, three repetitions per prompt, with every cited page fetched and scored.

  • Measured

    The evidence-graded citation report

    The Revenue Experts AI citation report labels every finding MEASURED or ADVISORY, so what we observed and what we recommend never blur. Per-platform precision reported separately with the spread shown. Citation stability across the three repetitions. Every cited URL listed with its support score, including the poor ones.

  • Advisory

    The AEO Improvement Plan

    The Revenue Experts AI improvement plan clusters gaps into page families, lists specific pages in priority order with what to change on each and why, and supplies a schema kit, an llms.txt and a sequenced roadmap. Built from what was on the pages we fetched, not from a generic checklist. The writing pipeline every resulting page goes through, including the write-first-verify-second rule, is published in full.

  • Included

    The AEO Visibility readiness assessment

    Citation readiness, content structure, authority signals, technical accessibility and semantic clarity. Sold separately at $29.90 per page. Included here, because a citation finding without a readiness finding tells you that you are not cited but not why.

  • Disclosed

    Our measurement disclosure

    The Revenue Experts AI measurement disclosure states engines and model versions, prompt count and construction, repetitions per prompt, how support is scored, where the two human gates sit, and which engines we exclude and on what grounds. Written against the eleven items the IAB names as required, including where we fall short.

$1,495

One-time price. No retainer, no annual commitment.

Published market pricing for a standalone AI visibility audit runs from $1,500 to $5,000, with mid-market retainers from $3,000 to $15,000 per month.Pricing

Book thirty minutes. We will run a handful of your category's prompts live on the call and open whatever comes back. If there is nothing worth auditing, we will say so on the call.

Book a call

Or test the argument before you test us

If you would rather check this yourself first, send your current provider three questions.

Do you fetch the cited page, or do you score from the answer text. How do you determine which claim a citation supports. How many times is each prompt run, and what is the variance.

It costs nothing and will tell you more about your current measurement than another month of dashboard access. If they answer all three well, keep them. That is a good outcome and you should be pleased. If they cannot, you now know what your number is worth, and you know it from the questions rather than from me.

The Revenue Signal

One shift a week, held to the same evidence standard as this page

If method-level thinking is useful to you, The Revenue Signal delivers one shift every Thursday: a real change in how B2B buyers research vendors, a named company that has already responded to it, and a concrete move you can make before the next issue lands.

Nothing goes in that cannot be traced to a named, dated source. Same rule as everything above.

Subscribe to The Revenue Signal

Every claim on this page, checkable

Sources

  1. IAB, Measuring Visibility in the AI Era. Published 3 August 2026, 36 pages, part of Project Eidos. All page references on this page are to this document.
    iab.com/guidelines/measuring-visibility-in-the-ai-era
  2. SparkToro. Rand Fishkin and Patrick O'Donnell, 27 January 2026. 600 volunteers, 12 prompts, 2,961 runs across ChatGPT, Claude and Google AI.
    sparktoro.com/blog
  3. Fractl, AI Search Consumer Trust Study. Q2 2026 survey of 1,008 US consumers and 150 marketers.
    frac.tl/ai-statistics
  4. Minuttia, State of AEO 2026.
    minuttia.com/state-of-aeo
  5. Semrush AI Visibility Index 2026. 126 million US AI search prompts, January to April 2026, across more than 1,200 brands and 22 verticals.
    ai-visibility-index.semrush.com · release announcement
  6. Similarweb, AI Citation Analysis. Citation volatility and the 11% ChatGPT to Perplexity overlap, according to Similarweb, sourced by Similarweb to an August 2025 GEO white paper.
    similarweb.com/blog/marketing/geo/ai-citation-analysis
  7. Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization, 2023 to 2026. Survey of 45 studies, arXiv preprint 2607.14035, July 2026.
    arxiv.org/abs/2607.14035
  8. SERP classification, "answer engine optimization services". All 98 ranked results, Google US, 10 July 2026, classified by an independent practice lab.
    geoplaybooks.com
  9. Published AEO and GEO pricing bands, 2026. Standalone audit $1,500 to $5,000, mid-market retainers $3,000 to $15,000 per month.
    pierview.ai pricing guide

Run figures quoted on this page come from our own reference run against revenueexperts.ai during pipeline validation, not from a client engagement. Pipeline behaviour described here reflects the Revenue Experts AI internal build and validation record.

Revenue Experts AI

We sell the AI Citation Audit, which means we compete in the market this page describes. Read it with that in mind. Our measurement disclosure is written against the IAB's required items, including where we fall short, so you can apply the same questions to us that we are asking you to apply to everyone else.

Findings on this page are derived in part from AI-generated responses across Claude, ChatGPT and Perplexity, with underlying source websites cited.

Book a call

Discover more from Revenue Experts AI

Subscribe now to keep reading and get access to the full archive.

Continue reading