How the OAIS run was done, and what an inconclusive result looks like when it is pre-registered

This is the reference for the second company case in our citation program. The story of the run is on LinkedIn, and the argument is on Medium. This page is the part you check.

What was fixed before the first query?

01-what-was-fixed-before-the-first-query

Everything that could bias a result was written down, hashed, and pushed to a version-controlled repository before any question was sent.

  • The prior. Predicted agreement between the mention ranking and the owned-citation ranking, Spearman’s rho, 0.70, range 0.55 to 0.85. Recorded 8 September 2026. A lean was recorded beside it: below 0.70 rather than above, with the reasoning. The lean carries no weight in the verdict; it is there to be scored across categories.
  • The decision rule. Rho at or above 0.85: the two rankings agree closely, and the public claim that citations reorder the picture is narrowed to exclude this category. Rho at or below 0.70: the claim is supported, for this category and this date only. Between: inconclusive. A point estimate within 0.02 of a threshold is reported as at the edge, never as having cleared it.
  • The guards. Search: every engine at least 90 percent complete and at most 3 of 36 answers without confirmed search. Sparsity: at least four ranked brands with own-page citations in three or more answers, or the verdict is “insufficient citation data”. Power: at least 10 ranked brands, or no band is read.
  • The questions. Twelve, one per slot of a fixed frame, generated to be brand-neutral, screened, six regenerated for slot fit with the rejected text logged, then frozen. SHA-256 ea22952e.
  • The brand list. Twenty brands, each with name variants and owned domains from the vendor’s own site and directories, never from an assistant’s answer. Common-word names on a review queue; “AI Shadow” matched only in that exact word order so that the phrase “shadow AI” can never count for it. Owned domains scoped to the product path for ibm.com, microsoft.com and servicenow.com. SHA-256 bce2c63c.
  • The engines. ChatGPT gpt-5.6-luna with the web search tool forced; Claude claude-opus-4-8 with the web search tool and a prompt instruction to search; Perplexity sonar-pro. Gemini excluded under Google Cloud’s terms. Settings export SHA-256 10083dbb.
  • Freeze commit 43c2eae, pushed 18:53 UTC, 8 September 2026. The three hashes were recomputed from that commit on 9 September and matched the record.

 

What the run produced

what-the-run-produced

108 planned calls, 108 completed, 108 with confirmed search, on 8 September 2026. Denominator for every count: 108.

The assistants named 23 vendors that were not on the frozen list. Under a rule written before the run, each was searched with the category noun, its official site located, and its domain and variants appended to a dated additions file before any statistic was computed. Two names one assistant produced, “KairoNull” and “Blade Labs”, have no site anywhere on the web as of 8 September and were logged as probable fabrications, not scored. Cloud platforms, SIEM tools and model providers named as components of a build-your-own stack were excluded as a class, and the class is named in the file.

Fifteen matches went to the review queue: twelve bare uses of “Purview” in Microsoft contexts, and one each of Archer, Galileo and Braintrust in vendor lists. All fifteen were accepted, with the reason recorded per row and the 24 occurrences of “shadow AI” listed by answer so the reader can see none was counted.

Eighteen brands were named in at least two questions and ranked. All three guards passed.

 

The result

the-result

Rho 0.768. Kendall’s tau-b 0.594. Prompt-level bootstrap interval, 2,000 draws, 0.010 to 0.891. Answer-level interval 0.423 to 0.857. Both intervals cross both thresholds.

Under the rule: inconclusive for this category. The point estimate sits inside the pre-registered range and is not within 0.02 of an edge. The recorded lean was not borne out.

Per engine: ChatGPT 0.882, Claude 0.494, Perplexity 0.623. A secondary prediction that the engine order would repeat the first category (ChatGPT, Claude, Perplexity) did not hold; the predicted order of citation volume (Perplexity 628, Claude 578, ChatGPT 203) did.

The subject, OAIS: named in 0 of 108 answers, own site cited in 1. This was the pre-registered expectation. OAIS was chosen as a baseline: a young vendor with a product page and a handful of posts, and no content written for retrieval by assistants, measured before any of that work exists. The prior states that it was expected to be absent or near-absent from both rankings and that the absence would be the finding, not a data failure. Below the inclusion threshold on both lists. In 61 answers a competitor’s own page was cited while OAIS was neither named nor cited.

 

What was amended after the run, and why is it visible?

what-was-amended-after-the-run

Pre-registration amendments are permitted, dated, and listed with the old text quoted. Three came after the run.

  • Top-3 tie rule. Two brands tied at citation rank 3.5 occupy positions three and four. The report first counted them as flips. The study owner ruled they are top-3 tied, neither in nor out. Four flips, two ties. No statistic changed.
  • A limitation not declared in advance. Inclusion was mention-based, so brands whose own pages were cited without the brand being named sit outside the ranked 18 even where cited more than any ranked brand: Kiteworks 23, Kognitos 13, Openlayer 9, Trussed 7, VerifyWise 6, WitnessAI 5, against 9 for the top ranked brand. The correlation is over engine-named brands only, and in this category that excludes the most-cited pages. The rule stands for this run; a change to “named in two questions or own page cited in three answers” is open for decision before the third category is frozen.
  • Provenance of the prior. The prior and decision rule were drafted with AI assistance in a chat, reviewed, adopted and dictated by the study owner the same day, before any data. That wording replaced an earlier one, and the earlier one is quoted in the amendment.

 

What the numbers may be read as

what-the-numbers-may-be-read-as

The run is Custom tier: three runs per question, twelve questions, no repeating schedule. Counts may be reported. Rates, stability figures and trends may not. “Cited in 1 of 108 answers” is a count. Nothing here is a percentage of buyers.

What the study records: a brand name appearing in an answer; a brand-owned page being cited; whether that page supports the sentence beside it, checked separately and, for this run, by an automated pass labelled as such. What it does not establish: that an assistant recommended a brand, that a buyer chose a competitor, or that either ranking predicts sales.

One category, one day, three assistants. The third category in the programme decides what may be said in public about mentions and citations; the rule is that the claim is revised only if two of three categories land in the same band.

Evidence package: https://drive.google.com/drive/folders/1_2lzly51d7i3v1vcopVgmnyOiTY5WSYU?usp=sharing It contains the pre-registration with every amendment, the frozen questions, the brand lists and additions, the engine settings, the freeze commit, the 108 derived answers, the page-check table and the report, and nothing about any individual.

To see the recommendations for OAIS, subscribe and read them on The Revenue Signal, with a self-check for your own category. Subscribe here

To have this run on your own category, book a fit conversation with John Bush: https://meetings.hubspot.com/john2956

Discover more from Revenue Experts AI

Subscribe now to keep reading and get access to the full archive.

Continue reading