How do you measure what an AI tool returned when you never recorded a baseline?

Title card reading, how do you measure what an AI tool returned when you never recorded a baseline, with the answer that it is rebuilt from records nobody cleaned: tickets, calendars, approvals, invoices and version history.

You rebuild the baseline after the fact, from operational records that were never meant to measure anything. Ticket systems, calendars, approval trails, supplier invoices and document version history all carry timestamps from before the tool existed, and none of them were deleted when it arrived.

The method below has six steps and takes a working day. It compares a reconstructed before period against the same calendar window one year earlier, names the rival explanations in writing before any figure is calculated, gives those rivals the benefit of every doubt, and allows three verdicts rather than two. The third verdict, cannot be determined, is the one most reviews are missing and the reason most of them produce fiction.

The Revenue Signal arrives every Thursday. One decision, one verified source, one move you can make before the next issue. Signing up delivers both the newsletter and the Executive AI ROI Scorecard. Get the scorecard and the newsletter, opens in a new tab.

Every guide to measuring AI opens with the same instruction. Establish a baseline before you deploy. It is correct, and it is useless to the person who needs it, because that person deployed eleven months ago and is now holding a renewal quote.

What follows is the method for that position. It is written for an executive to run, not a data science team, and every input it asks for is something a company of 50 to 5,000 people already has. Read it once to understand the shape, then work through it with the actual initiative in front of you.

Why is the vendor's dashboard not an answer?

Two columns comparing a number that can fail, which comes from records predating the tool and could have come out negative, against a number that cannot fail, which is usage multiplied by assumed minutes saved and returns the same answer if the tool is ignored.

A vendor dashboard reports two kinds of thing. The first is usage: sessions, queries, documents processed, seats active. Those are measured accurately and they are not returns.

The second is modelled savings: hours saved, cost avoided, productivity gained. These are calculated by multiplying usage by an assumed time saving per unit and an assumed hourly rate. Both assumptions were chosen by the party being paid.

The problem is not that these figures are dishonest. Most of them are produced in good faith. The problem is that they are unfalsifiable. There is no configuration of your business in which that dashboard reports a negative number, which means the dashboard cannot be evidence, because evidence is something that could have come out the other way.

Unfalsifiable is not the same as untrue. It does mean you cannot put it in a board pack and call it a result.

Where does a baseline you never recorded actually survive?

Five operational systems that still hold pre-deployment records: ticket and queue systems for cycle time, calendars for paid hours, approval trails for timestamps, supplier invoices for prior cost, and document version history.

Companies believe they have no pre-deployment data because they are looking in the analytics layer, where measurement is supposed to live. The records that survive are in the operational layer, where nobody was measuring anything and therefore nobody curated, aggregated or deleted.

SystemWhat it yields for the before period
Ticket, queue or shared inbox systemsVolume of work items, created and resolved timestamps, so cycle time and backlog are both recoverable to the day.
CalendarsThe recurring meeting that existed to do this work manually. Its frequency, duration and attendee list convert directly into hours of paid time.
Approval and workflow trailsSubmission and sign off timestamps in the ERP, the purchasing workflow or the e-signature tool. Elapsed time per approval, before and after.
Supplier and contractor invoicesThe agency, outsourcer or freelancer who did this work before. Monthly spend, and whether it stopped, shrank or continued after deployment.
Document and spreadsheet version historyWho touched the artefact, how often, and how long a cycle took, held automatically by the file system or the office suite.

Three of these five are usually enough. If none of them contains anything about the work in question, that is itself a finding, and it means the initiative was never attached to a repeatable process. That case is dealt with in step six.

What are the six steps?

The six steps numbered: name the task not the tool, set the windows one year apart, pull the metric from one source system, write the rival explanations down before calculating, cost it in full including maintenance and mediation, and choose one of three verdicts.

Step 1. Name the decision or task, not the tool

Write one sentence describing the recurring piece of work the tool touches. "Quoting a custom order." "Closing a support ticket." "Producing the monthly demand forecast." Not "the AI assistant."

This step exists because tools sprawl across several processes and a return can only be measured against one. If the tool touches four processes, you run this method four times or you pick the one that carries the most cost.

Step 2. Choose the comparison window, and choose it before you look at anything

Use the same calendar window one year before deployment. Not the weeks immediately preceding it.

This is the step most reviews get wrong, and getting it wrong changes the answer rather than the precision of the answer. Something goes wrong shortly before a company buys a tool. A backlog builds, a person leaves, a customer complains. That episode is frequently the reason the purchase was approved. Measuring against it compares the new tool against the worst period the process ever had, and produces an improvement that would have appeared anyway once the crisis passed.

Matching the calendar window also holds seasonality still, which matters in any business with a quarter end, a season or a renewal cycle.

Step 3. Write down the rival explanations before you calculate the difference

List everything other than the tool that could explain a change between those two windows. Do this in writing, and do it before you know what the difference is, because a list written afterwards is a list of things you have already decided to dismiss.

  • Headcount in the team went up or down
  • Volume of work went up or down
  • A process or policy change shipped in the same window
  • A systems change, such as a new ERP, CRM or ticketing migration
  • A personnel change, particularly a new manager of the process
  • Market or seasonal conditions that moved the work itself

Beside each one, write present or absent. For each one marked present, write in a sentence how much of a change it could plausibly account for on its own. You are pre-committing to an interpretation while you are still capable of being neutral about it. The technique is borrowed from research design and it is the single thing that separates a review from a story.

Step 4. Compute the difference, then hand the rivals every benefit of the doubt

Now take the before figure and the after figure for the same metric, from the same source, over matched windows.

Then subtract, generously, what the rivals you marked present could explain. Deliberately over-attribute to them. If headcount rose by one person in a team of six, assume the extra person explains a full sixth of any throughput gain before the tool gets credit for anything.

What survives that treatment is a floor, not an estimate. A floor is what you want, because a floor is the number that stays standing when the CFO attacks it.

Step 5. Build the cost side properly, which is where most of the error lives

Compare the surviving benefit against the full run cost, not the licence. Four lines, and the last two are the ones normally omitted.

  1. Licence or subscription, annualised.
  2. Implementation and integration, amortised over the realistic life of the tool rather than the contract term. If your category replaces itself every two years, amortise over two.
  3. Internal maintenance time. The hours your own people spend keeping it working, retraining it, fixing its inputs, checking its outputs.
  4. Mediation time. The hours a person spends running the tool on behalf of someone else and passing the answer along. This one is invisible in every budget and it is often the largest of the four.

Step 6. Choose one of three verdicts

Not two. Three.

  • Returned. The surviving benefit exceeds the full run cost. Record the floor figure and the rivals you subtracted, so the number can be defended in six months by someone who was not in the room.
  • Did not return. The full run cost exceeds the surviving benefit. This is a renewal decision, not a failure report.
  • Cannot be determined. The records do not exist, the windows are not comparable, or the rivals absorb the whole difference.

Cannot be determined is a legitimate result and it must be allowed, or the exercise becomes theatre. When two verdicts are permitted and the evidence supports neither, the review does not stop. It picks the one that suits the person presenting.

A cannot be determined verdict carries a required action, which is what stops it becoming a way of avoiding the question. Instrument the two numbers now, in the operational system where the work happens. Set a review date inside 90 days. Put the initiative on notice in writing, so the next renewal has evidence attached to it whichever way it goes.

What does this look like when it is finished?

The five things the finished one page verdict holds: the task and windows and metric, the raw difference with rivals marked, the surviving benefit, four cost lines, and the verdict with its follow up.

One page. The decision or task in a sentence. The two windows with their dates. The metric and its source system. The raw difference. The rivals marked present, with what each was assumed to explain. The surviving benefit. The four cost lines. The verdict, and if it is cannot be determined, the two numbers now being instrumented and the review date.

That page is the artefact. It is what you take into the renewal conversation, and it is what the next person to hold your job reads when they ask why this contract exists.

The Retrospective Baseline Worksheet, opens in a new tab lays the six steps out as a working page you can fill in and download. Nothing you enter is stored or sent anywhere.

What are the honest limits of this method?

Four stated limits: it is not a controlled experiment, it works badly on quality effects, it cannot see past the window, and it rewards work that was already legible.

Four, stated plainly, because a method that claims no limits is a sales document.

  • It is not a controlled experiment. There is no control group. Matching calendar windows and subtracting rivals reduces the error; it does not remove it. Treat the output as a defensible floor rather than a measurement.
  • It works badly on quality effects. If the tool improved accuracy, judgement or customer experience rather than time or cost, the operational residue will not hold it. Those need their own measurement and this method should not be stretched over them.
  • It cannot see effects longer than the window. Something that changes retention or renewal rates shows up over years, not quarters.
  • It rewards processes that were already legible. Work that never touched a ticket, a calendar or an invoice leaves no residue, and for that work the honest verdict is cannot be determined on the first pass.

Why do most AI reviews never reach a verdict at all?

Figure card reading 57 percent of enterprises say their AI returns still do not outpace their spend, unchanged since 2025, with the full survey scope and the note that the study did not ask whether they measured and found nothing or had no way to measure.

The aggregate picture is consistent with a lot of reviews that never produce an answer. Domino Data Lab published its Fifth Annual Enterprise AI Report on 21 July 2026. Across 639 senior AI leaders, the share of enterprises whose returns fail to outpace their spend held at 57 percent, unchanged from 2025, while 93 percent reported that their production capability improved over the same period.

Scope, stated before that figure is used any further. Vendor commissioned research: Domino Data Lab sells a platform in this category. Fielded independently by BARC Research in April 2026, among 639 leaders at director level and above, at organisations with annual revenue of 100 million dollars or more, across North America, the United Kingdom and continental Europe, in financial services and insurance, life sciences and public sector. Manufacturing, retail, consumer goods and energy technology were not surveyed. Every figure is self reported. Read it as direction, not decimal, and do not treat it as a measurement of your industry.

What that number cannot tell you is whether the 57 percent measured carefully and found nothing, or never had a way to measure at all. Both are consistent with the data, and the survey did not separate them. The method above is written for the second case.

Frequently asked questions

Can you measure AI return on investment without a baseline?

Yes, by reconstructing the baseline from operational records that predate the tool. Ticket systems, calendars, approval trails, supplier invoices and document version history all carry timestamps from before deployment. The reconstruction compares that period against the same calendar window one year earlier, subtracts what rival explanations could account for, and produces a defensible floor rather than a precise measurement.

Why compare against the same window one year earlier rather than the period just before deployment?

Because a problem usually occurs shortly before a company buys a tool, and that problem is often why the purchase was approved. Comparing against it measures the tool against the worst period the process ever had, which produces an improvement that would have appeared anyway once the episode passed. Matching the calendar window also holds seasonality constant.

Why is a vendor dashboard not evidence of return?

Because it reports usage accurately and savings from assumptions chosen by the party being paid. There is no configuration of a business in which such a dashboard reports a negative result, and a figure that could never come out the other way cannot function as evidence. That does not make it false. It makes it unusable as proof.

What should happen when the answer is that the return cannot be determined?

Cannot be determined is a legitimate verdict and should be one of three permitted outcomes rather than being resolved into a guess. It carries a required action: instrument the two relevant numbers in the operational system now, set a review date inside 90 days, and record in writing that the initiative is on notice so the next renewal decision has evidence attached to it.

Run this across the whole program instead of one initiative

The Executive AI ROI Scorecard applies the same logic across four areas of an AI program and returns a written result rather than a sales call. It takes about ten minutes and it tells you which initiatives can be assessed with the records you already hold and which cannot.

Take the Executive AI ROI Scorecard, opens in a new tab

The Revenue Signal is published every Thursday: one decision a senior executive faces in the next 90 days, one verified source, one move. Signing up delivers the newsletter and the Executive AI ROI Scorecard together. Sign up here, opens in a new tab.

Sources

  1. Domino Data Lab, "AI ROI Fails to Outpace Spend for 57% of Enterprises, Unchanged Since 2025, Even as 93% Now Report Improved Production," 21 July 2026. Supports the 57 percent and 93 percent figures, the 639 respondent count and the survey methodology. Verified live 22 August 2026. Read the source, opens in a new tab. Full report at domino.ai/enterprise-ai-report, opens in a new tab.

The six step method described in this article is Revenue Experts' own working procedure. It is not drawn from a published study, and it is presented as a method rather than as a finding.

Discover more from Revenue Experts AI

Subscribe now to keep reading and get access to the full archive.

Continue reading