Answer

Our AI pilot failed. What now?

Justas Butkus is a fractional AI officer based in Vilnius, Lithuania, working with mid-market companies and scale-ups across the UK, EU and US at two to four days a month. He builds and operates the production AI systems he advises on, and takes no commission, referral fee or revenue share from any vendor.

Short answer

A failed pilot is the statistically normal outcome. What matters now is the post-mortem, because the same decision that killed this one will kill the next one. In almost every case the cause is upstream of the technology: the wrong process was chosen, the data could not support it, or nobody would accept the risk of going live.

First: this is the normal outcome, not a personal failure

That is not consolation, it is context you will need for the conversation upstairs. RAND, which interviewed 65 data scientists and engineers with at least five years building production models, reports that by some estimates more than 80% of AI projects fail, roughly twice the rate of IT projects that do not involve AI. MIT NANDA found 95% of enterprise generative AI pilots delivering no measurable financial return.

The reason that matters to you specifically: Dataiku and the Harris Poll, surveying 900 CEOs across eight countries, found 80% of CEOs saying their job is at risk if AI fails, and 87% saying they would stake their job on delivering results from their AI initiatives. The pressure you are feeling is not unusual and it is not a signal that you handled this badly.

What is unusual is running a proper post-mortem instead of quietly starting another pilot. Almost nobody does it, which is why the second pilot usually fails for the same reason as the first.

The seven-question post-mortem

Work through these in order and stop at the first one where the honest answer is uncomfortable. That is your cause.

  1. Why this process, and who chose it?If the answer traces back to a demonstration somebody saw, you have found it. The justification was assembled after the fact to fit a solution that already existed. This is the single most common cause.
  2. What did one occurrence of that process actually cost, before we started?If nobody can answer, the project never had a business case, only a hypothesis. You cannot measure a return against a baseline you never took.
  3. Did the same field mean the same thing in every system it came from?Not whether the data existed. Whether it was consistent. Inconsistent definitions are learned faithfully by any model, and this is discovered late because everyone assumes their own data is fine.
  4. Who was going to accept the risk of putting it in front of customers?Name the person. If you cannot, the pilot was never going to ship regardless of how well it performed, and it became a permanent pilot by default.
  5. What number did we agree in advance would count as working?Agreed before the build, in writing. Without it, evaluation becomes an argument about impressions that nobody can win, and the project dies of ambiguity.
  6. Did anyone whose work it changed have a say before it was built?Systems that arrive as a surprise get worked around. This is not resistance to change, it is a design failure that was avoidable.
  7. Was this the first thing we ever put into production?If so, you attempted the hard part of the use case and the hard part of operating AI simultaneously. That combination produces most pilot graveyards.

Which failure is yours?

Diagnosing a failed AI pilot by symptom
What you observedActual causeWhat has to change
It worked and nobody used itWrong process chosenRe-select on cost and frequency, not on how well it demos
Good in testing, poor in realityInconsistent data definitionsFix definitions before touching a model
Endless further evaluation roundsNobody would accept go-live riskName an accountable owner and define the rollback
Legal or compliance stopped itGovernance considered too lateDesign oversight in from the start
Staff quietly worked around itNobody affected was consultedInvolve them before the build, not at training
Ran out of budget mid-buildScope was never boundedFixed scope with a defined end date
Vendor delivered what was asked, it was wrongSpecification error, not delivery errorThe specification is the work

What to tell the board

The instinct is to minimise. That is the wrong move, because a second failure after a minimised first one is much harder to survive than an honestly reported first one.

  • Lead with the cause, not the outcome. "We chose the wrong process and here is how we now choose" is a management answer. "It did not work out" is not.
  • Give the industry context once, briefly, and move on. More than 80% fail. Use it to establish that this is a known pattern, not to excuse anything.
  • Say what it cost and what was learned. Both numbers. The learning has to be specific enough to be checkable.
  • Bring the revised selection method. Not a new project. The method by which the next candidate will be chosen.
  • State the stopping rule. At what point you would stop the next one, decided in advance. Boards trust plans that can be abandoned far more than plans that cannot.

Should you try again, or stop?

A genuine answer, including the case for stopping.

  • Try again if you can now name a process that is frequent, expensive and unloved, cost one occurrence fully loaded, and multiply it out to an annual number that is obviously larger than the whole cost of the work.
  • Try again if the failure was ownership or data definition, both of which are fixable and cheaper to fix the second time.
  • Stop for now if no process survives that arithmetic. There is no shame in it and it is more common than the market admits.
  • Stop for now if the underlying process is broken. Automation will make a broken process faster and no better, and you will have paid to accelerate the problem.

Deciding to stop is a legitimate outcome of a post-mortem. It is also the outcome nobody selling you an AI project will ever recommend.

Frequently asked questions

How common is it for an AI pilot to fail?

Very. RAND reports that by some estimates more than 80% of AI projects fail, about twice the rate of IT projects that do not involve AI, and MIT NANDA found 95% of enterprise generative AI pilots delivering no measurable financial return. A failed pilot is the normal outcome rather than the exception.

Why do AI pilots fail?

Rarely because of the model. The recurring causes are choosing a process because it demonstrated well rather than because it was expensive, discovering late that data definitions differ across systems, and having nobody willing to accept the risk of putting the system in front of customers.

Should we try a different AI model or vendor?

Usually not. A different model solves a capability problem, and most failures are selection, data or ownership problems. If the process was not worth automating, a better model produces a better answer to a question nobody was asking.

What should we tell the board after a failed AI pilot?

Lead with the cause rather than the outcome, state what it cost and what was learned, present the revised method for selecting the next candidate rather than a new project, and define in advance the point at which you would stop.

How do we choose a better second project?

Rank candidates by annual frequency multiplied by fully loaded cost per occurrence, discount by the share genuinely automatable, and compare against the entire cost of specification, build and running. The right second project is usually boring and high-volume.

Is it reasonable to stop trying with AI?

Yes, and it is an underrated outcome. If no process survives the cost arithmetic, or if the underlying process is broken, stopping is correct. Automation applied to a broken process makes it faster and no better.

If you are writing the post-mortem now

The seven questions above are the whole method. If you want a second opinion on what they surfaced, that is a short conversation.