The Revenue Experts AI Citation Audit Method
How we measure whether B2B SaaS companies get cited by ChatGPT, Claude, Perplexity, and Gemini.
Why this method exists
The B2B AI search visibility space has no standard for measuring who gets cited and why. Agencies report "share of voice" against vague metrics. Buyers cannot tell whether a $10,000-a-month engagement moved the needle or the agency rebranded SEO work and changed nothing.
The shift in how B2B buyers research vendors makes this measurement gap expensive. 73% of B2B buyers now use AI tools like ChatGPT and Perplexity in their research process[1]. AI search traffic converts at 14.2% versus Google organic's 2.8% — a 5.1x advantage[1]. 51% of B2B software buyers now start their research in an AI chatbot rather than Google[2]. If your buyers are starting in AI and you don't know whether AI cites you, you can't manage the channel. (We covered the underlying problem in detail in B2B AI Visibility: Why Your Website Is Invisible to ChatGPT.)
We built this method to fix that. It produces three things any B2B SaaS founder should be able to ask for and get: a number that says where you stand, a list of who beats you and why, and a sequenced fix list with effort and expected impact attached.
The method runs the same way every time. Same prompt, LLM and scoring. Repeatability matters because re-running the audit in 90 days produces a comparable score, not a different methodology that happens to agree with the first one.
What the method tests
Three things, in plain language:
- Whether your company appears in AI search results for the queries your buyers actually ask
- Which competitors get cited when you don't
- What kind of content gets cited (so you know what to build, not just what to fix)
That's it. The method does not measure brand awareness, ad recall, organic Google rankings, or social engagement. Those matter, but they are not the same problem. AI citation is its own measurement problem with its own answer. For our broader services connected to AI revenue (custom AI specialists, marketing automation, revenue intelligence), see the services overview.
The 50 prompts and their four categories
Every audit runs 50 prompts. They are divided into four categories that match how B2B SaaS buyers actually search.
Commercial intent
Vendor-evaluation queries. "Best AEO consultants for B2B SaaS." "Best RAG implementation companies." "Top AI consultants for Series B startups." When buyers shop the category, these are the queries they type.
Transactional intent
Cost, timing, scope queries. "How much does an AI visibility audit cost?" "Custom RAG implementation pricing." "How long does an AEO sprint take?" These appear when buyers move from awareness to budget.
Comparison intent
Direct vendor-vs-vendor queries. "Vstorm vs Stratagem for RAG." "Discovered Labs alternatives." "Animalz vs Embarque for B2B SaaS." Late-stage buyers compare specific options.
Decision intent
Build-vs-buy and hire-vs-consultant questions. "Should we hire an AI engineer or run a sprint?" "Custom RAG or platform?" "Fractional advisor or full-time?" These are the questions buyers ask after they have shortlisted but before they commit.
The exact 50 prompts live in our internal scoring system. We do not publish them. Publishing them would make the method reproducible by competitors who have not tested it, which lowers its value to clients who pay us to run it.
What we do publish: the categories, the count per category, the source rationale for prompt selection, and the methodology for adapting prompts to a specific company's industry segment.
The four LLMs we test
ChatGPT (GPT-4 / GPT-4 Turbo via OpenAI). Claude (Sonnet / Opus). Perplexity (Sonar models with web search). Gemini (Pro with Google search integration).
We chose these four for reasons that aren't obvious if you only use one of them:
- ChatGPT favors blog content from established domains. Wikipedia accounts for roughly 47.9% of its top citations[3]. It cites less often than the others but with stronger authority weight per citation.
- Claude pulls from professional, well-structured sources. It rarely cites Reddit. It cites methodology pages and white papers more than competitor LLMs.
- Perplexity cites Reddit heavily. Earlier research from Bluefish AI found Reddit accounted for 46.7% of Perplexity's top-10 citations[3], though that share dropped to roughly 24% by January 2026 after Reddit's lawsuit against Perplexity in October 2025[4]. Perplexity surfaces newer content faster than the others.
- Gemini integrates Google search behavior. Its citations correlate with Google AI Overviews more than the other three. ChatGPT leans on listings for 48.7% of local source citations, while Gemini favors websites at 52.1%[1].
Different LLMs surface different competitors. A company can be cited well on ChatGPT and invisible on Perplexity, or owned on Perplexity and absent on Claude. Testing across all four is what reveals the actual citation map.
Citation classification
Not every "appearance" is a citation. The method classifies every result four ways.
Source type
Five buckets: your domain, your domain mentioned by name without link, competitor cited (with link or named), aggregator cited (G2, Capterra, listicle blogs), or no relevant citation at all. The first two are wins. The third is a loss. The fourth is partial. The fifth is open territory.
Authority weight
First-position citations (named in the answer summary or top of the cited list) score higher than late-position citations (named only in expanded results or "see also" lists).
Reproducibility
Each prompt runs three times across each LLM, so 12 runs per prompt and 600 runs per audit. Citations that appear consistently across runs score higher than citations that appear once and don't reproduce. LLM responses vary. The method accounts for that.
Linked vs unlinked mentions
A page that mentions Revenue Experts AI without linking still helps. The LLM associates the entity with the topic. Linked citations score higher, but unlinked mentions count. The strongest single predictor of whether a brand appears in AI answers is how frequently that brand is mentioned across authoritative web sources — regardless of whether those mentions include a hyperlink[1].
The five outcome buckets
After every audit, every prompt lands in one of five buckets:
- Cited and dominant — your firm appears in at least three of four LLMs, in first or second position. This is the goal state.
- Cited but weak — your firm appears in one or two LLMs, in late position or as one of many. Defensible but not winning.
- Open territory — no specialist firm consistently cited across the four LLMs. Good content can take this position.
- Competitor-owned — one or two named competitors dominate. Direct fight required to displace them.
- Saturated — five or more competitors cited. Insertion is possible. Dominance is unrealistic in 12 months.
The bucket distribution across your 50 prompts tells you where to invest content effort. Open territory prompts are quick wins. Competitor-owned prompts need direct comparison content. Saturated prompts are usually deprioritized in the first 90 days.
Domain frequency analysis
For every audit, we count which domains get cited across all 600 runs. The output is a ranked list of cited sources for the buyer's category.
This is one of the most useful artifacts the audit produces, because the cited domain list is rarely what the company expected. A B2B SaaS that thought it competed with Vendor X often discovers that the LLMs cite Vendor Y, Vendor Z, and a Reddit thread it never knew existed. The competitor map is different from the Google-search competitor map. Independent research found the top 15 domains absorb 68% of the AI answer pipeline across all categories[5].
The domain frequency list also tells you what content to build. If LoudFace's listicle ranks third on three of four LLMs, you study its structure. If a Reddit thread by a specific user dominates Perplexity citations for your category, that user's contribution pattern becomes your model for Reddit engagement. Our Competitive Intelligence as a Service tracks competitor citation patterns continuously between audits.
What this method finds at the category level
Some patterns repeat across most audits we run for B2B SaaS companies:
- Most unbranded commercial queries cite zero specialist B2B AI consulting firms by name in the first answer position. Most LLMs default to general guidance or aggregator results before naming specific firms.
- The same 12-15 firms dominate cited space across most B2B AI consulting categories. Citation concentration is high — the top 15 domains absorb 68% of the AI answer pipeline across all categories[5].
- Pages with HTML pricing tables are extracted at higher rates than pages with image-based pricing. LLMs treat HTML tables almost as direct quotes.
- Company pages with named-client outcome strings (a specific result like "we grew Client X from N to N+M in W weeks") appear in citation positions far more often than pages without specific outcome data. Independent audit research found the top quartile of cited SaaS pages get cited 8.4× more often than the bottom half — and the gap tracks to a small set of structural choices, not domain authority[6].
- Domain age and existing authority help but are not the deciding factor. Newer domains with strong topical content can beat older domains with thin content[6]. We catalogued these patterns in detail in The 36 AI Search Visibility Factors.
- Each LLM rewards different content patterns: ChatGPT rewards comparisons, Perplexity rewards depth, Claude rewards methodology, Gemini rewards schema[6].
These are category-level observations, not your company's specific findings. The audit produces your company's specific findings.
How this method differs from existing alternatives
A few B2B AI agencies and tools have published their own methodologies. The relevant differences:
DerivateX's approach
Tests 1,400 prompts across 50 B2B SaaS companies. Bigger sample, less depth per company. It is better suited for a category report than a per-company audit. Our method is per-company at higher diagnostic depth.
Share-of-voice tools (LoudFace, Otterly, Profound, AthenaHQ)
They estimate citation share through modeled metrics and limited prompt samples. Faster but less specific about what to fix. Our method runs actual citations on actual prompts, which means the output is "you appeared at position 3 on Perplexity for Prompt 17" rather than a percentage estimate.
Gartner Magic Quadrant equivalents
They measure analyst perception. They matter for procurement but don't reflect what AI engines actually cite today. Our method measures live citation behavior, not analyst opinion.
None of these are wrong. They measure different things. We built this method because we needed something more specific than share-of-voice and faster than a Gartner equivalent.
What this method doesn't measure
The method has real limits. We name them up front because the alternative is overpromising:
- It does not measure conversion. Citation is upstream of revenue, not equivalent. Some companies are cited heavily and convert poorly. The method tells you the visibility problem, not the conversion problem.
- It does not predict 6-month citation behavior. LLM citation patterns shift as models update. Citation graphs can shift faster than content strategies — Perplexity's Reddit citation share dropped 86% almost overnight after Reddit sued Perplexity in October 2025[4]. We recommend re-running the audit every 90 days for fast-moving categories.
- It does not account for personalized AI responses. Each user's history affects what they see. The method tests anonymized prompts. Your buyers' personalized results may differ.
- A 50-prompt sample is enough for category-level findings, not statistically significant per-prompt findings. We don't claim p-values. We claim diagnostic clarity.
- The method can't measure citations on platforms we don't test (Microsoft Copilot, Bing Chat, smaller AI search engines). We add LLMs as they reach meaningful market share.
If a tool claims to measure things our method doesn't, it's worth asking how. Some claims are real. Some are marketing.
Apply this method to your company
We run this method as our $497 AI Visibility Audit.
What you get:
- Your numerical citation score across the 50 prompts × 4 LLMs
- Your domain frequency map (who gets cited when you don't)
- Your bucket distribution across the 5 outcome buckets
- Specific gaps mapped to specific fixes
- A sequenced action list with effort, expected impact, and dependency order
- A 30-minute walkthrough call to explain results
Turnaround: 5-7 business days from kickoff.
Pricing
| Engagement | Price | When it fits |
|---|---|---|
| AI Visibility Audit | $497 | You want to know where you stand and what to fix. |
| Targeted fix | $997 - $1,497 | One specific gap is the priority (e.g., your pricing pages aren't extractable, or you have no comparison content). Fix it cleanly. |
| Foundation sprint (AI Revenue Blueprint) | $2,497 | Multiple gaps need to be fixed in sequence. Six-week scope, full content + technical foundation. |
You can also stop after the audit. Plenty of clients run the audit, get the action list, and execute it themselves. The audit is useful even if you never hire us for follow-on work. 100% of the audit fee credits toward the Foundation sprint if you decide to move forward.
Want to see the method applied to your company?
Book the $497 AI Visibility Audit. 5-7 day turnaround.
Book the audit →Frequently asked questions
How long does the audit take from start to finish?
5-7 business days. Day 1: kickoff call to confirm prompts match your industry segment. Days 2-5: we run the 600 LLM calls and classify results. Day 6-7: written report and 30-minute walkthrough.
What format is the audit delivered in?
A 12-15 page PDF report and a spreadsheet with the per-prompt data. The spreadsheet is exportable so your team can re-test specific prompts later.
How is this different from an SEO audit?
SEO audits measure rankings on Google search. This audit measures citation appearance on AI engines. Different mechanism, different scoring, different fixes. A site can rank well on Google and not get cited by ChatGPT. The opposite happens too.
Do you guarantee citation improvements?
No. Anyone who guarantees AI citation improvement either does not understand how LLMs cite content or is selling something else. We guarantee the audit will tell you what's measurable, what's broken, and what to fix in priority order. Whether you act on it determines the outcome.
How often should I re-run the audit?
Every 90 days for fast-moving categories (B2B AI consulting, AEO, RAG implementation). Every 6 months for slower-moving categories. Re-running uses the same 50 prompts so the score is comparable across audits.
Can my SEO agency run this for me?
They can if they want to. The method is published. The 50 prompts are ours. SEO agencies generally don't run AI citation audits because the scoring system is different, the source classification is different, and the fix list is different. If your SEO agency wants to add this capability, we'd rather they audit your site themselves than misrepresent SEO work as AEO work.
What if my company is too small for this?
The audit works for any B2B SaaS company past product-market fit. If your monthly recurring revenue is under $50,000, the audit is probably premature — your problem is sales, not visibility. Fix the upstream problem first.
Is the method specific to B2B SaaS?
Yes. The 50 prompts, the buckets, and the scoring assume B2B buyer behavior. We've tested it on B2B AI consulting, RAG implementation, AEO services, and adjacent categories. We haven't validated it for D2C, e-commerce, or non-B2B contexts. If you're outside B2B SaaS, ask before buying.
Method authors
Co-founder + Fractional AI Search Advisor
Runs the methodology, prompt library, and reporting structure. Owns everything connected to AI at Revenue Experts AI.
Co-founder + Chief Executive Officer
Runs business operations, client engagement strategy, implementation roadmap — how custom AI revenue specialists, agent architectures, and library integrations come together into working systems for B2B clients.
Elizabeta built the original 50-prompt set, the bucket framework, and the scoring rubric. She is a B2B strategist with 25 years of experience across digital marketing, finance, and technology, with an MBA in Finance and Marketing. She works as a Fractional AI Search Advisor for B2B SaaS companies in the Series A and Series B stage.
John leads the company as CEO. He oversees client relationships, business strategy, and the full delivery system that wraps around the audit. On the AI side, he architects how findings from the audit translate into custom AI specialists, RAG implementations, and revenue intelligence systems for clients who move into the build phase.
We run every audit ourselves. We don't subcontract the analysis. The methodology is ours, the runs are ours, the reporting is ours.
Sources and references
- Loganix, "2026 B2B AI Buying Behavior Analysis" (April 2026). Synthesizes six independent studies covering 680 million AI citations, 2,961 research sessions, and 1.96 million browsing sessions. Source for: 73% of B2B buyers using AI tools, 14.2% AI conversion vs 2.8% Google organic, ChatGPT 48.7% local listings, Gemini 52.1% websites, brand mention as the strongest citation predictor.
- G2, "The Answer Economy: G2's 2026 AI Search Insight Report" (March 2026). Survey of 1,076 B2B decision-makers across North America, EMEA, and APAC. Source for: 51% of B2B software buyers starting research in AI chatbot, 87% saying AI chatbots are changing how they research software, 69% choosing different vendor than expected because of AI guidance, 47% preferring ChatGPT.
- Bluefish AI / Discovered Labs, "AI Citation Patterns: How ChatGPT, Claude, and Perplexity Choose Sources" (December 2025). Source for: ChatGPT favoring Wikipedia (47.9% of top citations), Perplexity prioritizing Reddit (46.7% of top-10 citations historically). See also Profound, "AI Platform Citation Patterns" for citation volume analysis (August 2024 to June 2025).
- CMSWire, "Reddit's Rise in AI Citations: What Marketers Must Know About AEO Strategy" (March 2026). Citing Tinuiti's Q1 2026 AI Citation Trends report and Conductor research. Source for: Perplexity Reddit citations dropping to ~24% by January 2026, 86% drop after Reddit's October 2025 lawsuit against Perplexity, citation graphs shifting faster than content strategies.
- 5W Public Relations, "AI Platform Citation Source Index 2026" (May 2026). Synthesizes 680 million citations across ChatGPT, Google AI Overviews, Perplexity, Gemini, and Claude. Source for: top 15 domains absorbing 68% of the AI answer pipeline.
- Digital Applied, "500 SaaS Sites Audited: AI Citation Visibility 2026" (April 2026). Stratified audit of 500 SaaS landing pages across ChatGPT, Perplexity, Claude, and Gemini over a 30-day window. Source for: top quartile cited 8.4× more often than bottom half, structural choices outweighing domain authority, engine-specific weighting (ChatGPT rewards comparisons, Perplexity rewards depth, Claude rewards methodology, Gemini rewards schema).
Related
Book the audit
- Book the $497 AI Visibility Audit — 5-7 day turnaround
Read more on AI search visibility
- B2B AI Visibility: Why Your Website Is Invisible to ChatGPT (And How to Fix It)
- The 36 AI Search Visibility Factors That Determine Whether AI Systems Cite Your Website
- The 1,300% AI Search Traffic Surge: Why Your Business Can't Ignore AI Search Visibility
- Why Product Leaders Should Build Their Own MVP: Lessons from an AI Search Visibility Tool
Explore our services
- AEO Get Seen — our flagship AI search visibility service
- All AI services — custom AI specialists, agents, libraries
- Read more on the blog
Stay current
- Subscribe to The Revenue Signal newsletter — every Thursday