AI Strategy · 6 min read

The Missing Baseline: Why AI Programs Can’t Prove Progress, and Stall Because of It

Many AI programs stall without failing: nobody recorded how the work ran before the tools arrived, so improvement can't be shown. The missing baseline, not the model, is what ends the funding.

TL;DR

  • Many AI programs stall without any visible failure. They stall because nobody captured how the work ran before the tools arrived, so improvement can be felt but not demonstrated.
  • The missing baseline turns the first budget review into an argument about impressions. Impressions lose to line items.
  • Organizations that move past early adoption tend to treat the “before” state as an asset, recorded deliberately and early, while it is still observable.
  • A baseline is also the only reliable way to tell a real improvement from a shift in where the work, and the errors, now sit.

The stall that looks like success

The familiar picture of stalled AI adoption involves a failed pilot, a disappointing model, or staff who refuse to use the tools. A quieter pattern is more common and harder to diagnose. The tools are in use. People say they help. Nothing is visibly broken. And yet the program never graduates from experiment to infrastructure, and eventually its budget is folded into something else.

Analysis of this pattern points less to the technology than to the evidence the organization holds about itself. When the question “is this working?” arrives, usually from finance and usually during planning season, the team has testimonials and no measurements. The testimonials are sincere. They are also unfalsifiable, and an unfalsifiable claim cannot survive contact with a competing request that comes with a number attached.

This is distinct from the argument that adoption stalls when nobody owns the cause of errors. Ownership of corrections is a problem about the future behavior of the system. The baseline problem is about the past, and the past is the one thing that cannot be recovered after the fact.

Why the “before” state vanishes

Work that has always been done a certain way is rarely measured, because it was never in question. A regional distributor’s order-entry team knows the work is slow and error-prone in a general sense. It does not know how long a typical order takes, how often one is corrected downstream, or how many orders wait in an inbox on a Monday. Those quantities were absorbed into the working week.

Then a tool is introduced. Enthusiasm is highest in the first weeks, the same weeks in which measurement would have been cheapest. The team is busy learning, the sponsor wants momentum, and recording the old process feels like delay. Within a month the old process no longer exists in a form anyone can observe. People have adapted their habits around the tool, the inbox has been reorganized, and the memory of the earlier workload has softened into a vague sense that things used to be worse.

Three features of this make the loss structural rather than careless:

  • The cost of capture is front-loaded. Measuring the old process takes effort before any benefit is visible.
  • The benefit of capture is deferred. It pays off only at a later review that nobody is thinking about yet.
  • The window is closing as the work begins. Every week of tool use contaminates the comparison.

An organization acting rationally on the information available at launch will skip the baseline nearly every time. The incentives are arranged to produce the gap.

What the absence does to the first review

When the review arrives, three things tend to happen in sequence.

First, the team reaches for the available proxies: usage counts, number of prompts, hours of access, licenses assigned. These measure activity, not effect, and they climb steadily regardless of whether the work improved. Usage is the easiest number to produce and the weakest evidence of value, a distinction that parallels the broader gap between productivity and understanding.

Second, a skeptical reviewer asks the obvious question: compared to what? The team answers with recollection. Recollection is contested, because the people who liked the tool remember a slow past and the people who did not remember an adequate one. Neither side can be checked.

Third, the decision defaults to whichever option has clearer accounting. A competing initiative with a before-and-after figure, even a modest one, beats an AI program with an enthusiastic but unmeasured story. The program is not rejected. It is simply starved of the sponsorship required to move from a tool people use to a layer the organization depends on.

Seen through the three-stage framing of maturity, this is a specific mechanism for staying in the first stage. Moving beyond ad hoc use requires a commitment of structure and attention, and that commitment is granted on evidence. Without a baseline there is no evidence, so the commitment is never made, and the organization concludes that the technology had a ceiling when in fact the record-keeping did.

A baseline is not a benchmark

The remedy is easy to mistake for something heavier than it is. A baseline in this sense is not an industry benchmark and not a formal study. It is a plain description of the work as it currently runs, written down before it changes:

Element recorded What it later allows
Typical duration of a unit of work A comparison of effort, not sentiment
Where work waits and for how long Detection of bottlenecks that merely move
Rate and type of corrections Distinguishing fewer errors from relocated errors
Who touches the work and at which step Seeing whether handoffs increased

None of this requires sophisticated instrumentation. A few weeks of rough observation by the people doing the work produces a record that is crude and still far better than none. The standard it must meet is not precision. It is that someone other than the enthusiasts can read it and see what was true before.

Note what is absent from the table: any measure of the tool itself. The baseline describes the work, not the technology. This is deliberate. A baseline tied to a specific tool becomes obsolete when the tool changes, and in a field where the model stack is replaced every few quarters, tool-specific measurement has a short shelf life. A description of the work outlives several generations of tooling.

Relocated effort, the second reason to look back

The baseline does a second job that is easy to overlook. It reveals whether effort has been removed or merely moved.

A common observation in automated processes is that the step the tool performs gets faster while the surrounding steps absorb new work: checking output, reformatting it, reconciling it with other records. The visible task shrinks, and the invisible tasks grow. This is the same dynamic described in analyses of failures at process handoffs, seen from the measurement side. Without a record of the earlier workload, an organization cannot tell whether total effort fell or whether it was redistributed to people who were not counted.

Teams in this situation often report feeling busier after adoption despite faster individual tasks. That is not a contradiction. It is what relocated effort feels like from the inside, and only a before-state can confirm it.

What separates organizations that move past this

The organizations that progress share a habit that looks unglamorous. They record the condition of a process before changing it, treat that record as a controlled document, and revisit it when the review comes. They tend to define the comparison in advance, so the judgment cannot be bent afterward to fit the mood of the room.

The pattern suggests a reframing of maturity itself. Maturity is commonly described in terms of capability: more tools, deeper integration, more autonomy. A complementary reading is that maturity is the degree to which an organization can state, with evidence, what its AI use has changed. By that definition, an organization with modest tooling and a clean baseline is further along than one with sophisticated tooling and no record of the starting point.

The implication is uncomfortable for programs already in motion. The window for capturing the original state has closed for work already changed. What remains is to capture the current state now, as the baseline for the next change, and to accept that the earlier gains may never be provable. That is a real loss, and it is the usual price of treating measurement as something to do later.

Research questions on this pattern are welcome through Inquiries.

Want help applying this?

If this sounds like your organization, tell us what's going on. We usually reply within two business days with a suggestion on where to start.

Get in touch