You rebuild the baseline after the fact, from operational records that were never meant to measure anything. Ticket systems, calendars, approval trails, supplier invoices and document version history all carry timestamps from before the tool existed, and none of them were deleted when it arrived.
The method below has six steps and takes a working day. It compares a reconstructed before period against the same calendar window one year earlier, names the rival explanations in writing before any figure is calculated, gives those rivals the benefit of every doubt, and allows three verdicts rather than two. The third verdict, cannot be determined, is the one most reviews are missing and the reason most of them produce fiction.
The Revenue Signal arrives every Thursday. One decision, one verified source, one move you can make before the next issue. Signing up delivers both the newsletter and the Executive AI ROI Scorecard. Get the scorecard and the newsletter, opens in a new tab.
Every guide to measuring AI opens with the same instruction. Establish a baseline before you deploy. It is correct, and it is useless to the person who needs it, because that person deployed eleven months ago and is now holding a renewal quote.
What follows is the method for that position. It is written for an executive to run, not a data science team, and every input it asks for is something a company of 50 to 5,000 people already has. Read it once to understand the shape, then work through it with the actual initiative in front of you.
A vendor dashboard reports two kinds of thing. The first is usage: sessions, queries, documents processed, seats active. Those are measured accurately and they are not returns.
The second is modelled savings: hours saved, cost avoided, productivity gained. These are calculated by multiplying usage by an assumed time saving per unit and an assumed hourly rate. Both assumptions were chosen by the party being paid.
The problem is not that these figures are dishonest. Most of them are produced in good faith. The problem is that they are unfalsifiable. There is no configuration of your business in which that dashboard reports a negative number, which means the dashboard cannot be evidence, because evidence is something that could have come out the other way.
Unfalsifiable is not the same as untrue. It does mean you cannot put it in a board pack and call it a result.
Companies believe they have no pre-deployment data because they are looking in the analytics layer, where measurement is supposed to live. The records that survive are in the operational layer, where nobody was measuring anything and therefore nobody curated, aggregated or deleted.
| System | What it yields for the before period |
|---|---|
| Ticket, queue or shared inbox systems | Volume of work items, created and resolved timestamps, so cycle time and backlog are both recoverable to the day. |
| Calendars | The recurring meeting that existed to do this work manually. Its frequency, duration and attendee list convert directly into hours of paid time. |
| Approval and workflow trails | Submission and sign off timestamps in the ERP, the purchasing workflow or the e-signature tool. Elapsed time per approval, before and after. |
| Supplier and contractor invoices | The agency, outsourcer or freelancer who did this work before. Monthly spend, and whether it stopped, shrank or continued after deployment. |
| Document and spreadsheet version history | Who touched the artefact, how often, and how long a cycle took, held automatically by the file system or the office suite. |
Three of these five are usually enough. If none of them contains anything about the work in question, that is itself a finding, and it means the initiative was never attached to a repeatable process. That case is dealt with in step six.
Write one sentence describing the recurring piece of work the tool touches. "Quoting a custom order." "Closing a support ticket." "Producing the monthly demand forecast." Not "the AI assistant."
This step exists because tools sprawl across several processes and a return can only be measured against one. If the tool touches four processes, you run this method four times or you pick the one that carries the most cost.
Use the same calendar window one year before deployment. Not the weeks immediately preceding it.
This is the step most reviews get wrong, and getting it wrong changes the answer rather than the precision of the answer. Something goes wrong shortly before a company buys a tool. A backlog builds, a person leaves, a customer complains. That episode is frequently the reason the purchase was approved. Measuring against it compares the new tool against the worst period the process ever had, and produces an improvement that would have appeared anyway once the crisis passed.
Matching the calendar window also holds seasonality still, which matters in any business with a quarter end, a season or a renewal cycle.
List everything other than the tool that could explain a change between those two windows. Do this in writing, and do it before you know what the difference is, because a list written afterwards is a list of things you have already decided to dismiss.
Beside each one, write present or absent. For each one marked present, write in a sentence how much of a change it could plausibly account for on its own. You are pre-committing to an interpretation while you are still capable of being neutral about it. The technique is borrowed from research design and it is the single thing that separates a review from a story.
Now take the before figure and the after figure for the same metric, from the same source, over matched windows.
Then subtract, generously, what the rivals you marked present could explain. Deliberately over-attribute to them. If headcount rose by one person in a team of six, assume the extra person explains a full sixth of any throughput gain before the tool gets credit for anything.
What survives that treatment is a floor, not an estimate. A floor is what you want, because a floor is the number that stays standing when the CFO attacks it.
Compare the surviving benefit against the full run cost, not the licence. Four lines, and the last two are the ones normally omitted.
Not two. Three.
Cannot be determined is a legitimate result and it must be allowed, or the exercise becomes theatre. When two verdicts are permitted and the evidence supports neither, the review does not stop. It picks the one that suits the person presenting.
A cannot be determined verdict carries a required action, which is what stops it becoming a way of avoiding the question. Instrument the two numbers now, in the operational system where the work happens. Set a review date inside 90 days. Put the initiative on notice in writing, so the next renewal has evidence attached to it whichever way it goes.
One page. The decision or task in a sentence. The two windows with their dates. The metric and its source system. The raw difference. The rivals marked present, with what each was assumed to explain. The surviving benefit. The four cost lines. The verdict, and if it is cannot be determined, the two numbers now being instrumented and the review date.
That page is the artefact. It is what you take into the renewal conversation, and it is what the next person to hold your job reads when they ask why this contract exists.
The Retrospective Baseline Worksheet, opens in a new tab lays the six steps out as a working page you can fill in and download. Nothing you enter is stored or sent anywhere.
Four, stated plainly, because a method that claims no limits is a sales document.
The aggregate picture is consistent with a lot of reviews that never produce an answer. Domino Data Lab published its Fifth Annual Enterprise AI Report on 21 July 2026. Across 639 senior AI leaders, the share of enterprises whose returns fail to outpace their spend held at 57 percent, unchanged from 2025, while 93 percent reported that their production capability improved over the same period.
Scope, stated before that figure is used any further. Vendor commissioned research: Domino Data Lab sells a platform in this category. Fielded independently by BARC Research in April 2026, among 639 leaders at director level and above, at organisations with annual revenue of 100 million dollars or more, across North America, the United Kingdom and continental Europe, in financial services and insurance, life sciences and public sector. Manufacturing, retail, consumer goods and energy technology were not surveyed. Every figure is self reported. Read it as direction, not decimal, and do not treat it as a measurement of your industry.
What that number cannot tell you is whether the 57 percent measured carefully and found nothing, or never had a way to measure at all. Both are consistent with the data, and the survey did not separate them. The method above is written for the second case.
Yes, by reconstructing the baseline from operational records that predate the tool. Ticket systems, calendars, approval trails, supplier invoices and document version history all carry timestamps from before deployment. The reconstruction compares that period against the same calendar window one year earlier, subtracts what rival explanations could account for, and produces a defensible floor rather than a precise measurement.
Because a problem usually occurs shortly before a company buys a tool, and that problem is often why the purchase was approved. Comparing against it measures the tool against the worst period the process ever had, which produces an improvement that would have appeared anyway once the episode passed. Matching the calendar window also holds seasonality constant.
Because it reports usage accurately and savings from assumptions chosen by the party being paid. There is no configuration of a business in which such a dashboard reports a negative result, and a figure that could never come out the other way cannot function as evidence. That does not make it false. It makes it unusable as proof.
Cannot be determined is a legitimate verdict and should be one of three permitted outcomes rather than being resolved into a guess. It carries a required action: instrument the two relevant numbers in the operational system now, set a review date inside 90 days, and record in writing that the initiative is on notice so the next renewal decision has evidence attached to it.
The Executive AI ROI Scorecard applies the same logic across four areas of an AI program and returns a written result rather than a sales call. It takes about ten minutes and it tells you which initiatives can be assessed with the records you already hold and which cannot.
The Revenue Signal is published every Thursday: one decision a senior executive faces in the next 90 days, one verified source, one move. Signing up delivers the newsletter and the Executive AI ROI Scorecard together. Sign up here, opens in a new tab.
The six step method described in this article is Revenue Experts' own working procedure. It is not drawn from a published study, and it is presented as a method rather than as a finding.
Subscribe now to keep reading and get access to the full archive.