Answer
Our AI pilot failed. What now?
Last updated
Justas Butkus is a fractional AI officer based in Vilnius, Lithuania, working with mid-market companies and scale-ups across the UK, EU and US at two to four days a month. He builds and operates the production AI systems he advises on, and takes no commission, referral fee or revenue share from any vendor.
Short answer
A failed pilot is the statistically normal outcome. What matters now is the post-mortem, because the same decision that killed this one will kill the next one. In almost every case the cause is upstream of the technology: the wrong process was chosen, the data could not support it, or nobody would accept the risk of going live.
First: this is the normal outcome, not a personal failure
That is not consolation, it is context you will need for the conversation upstairs. RAND, which interviewed 65 data scientists and engineers with at least five years building production models, reports that by some estimates more than 80% of AI projects fail, roughly twice the rate of IT projects that do not involve AI. MIT NANDA found 95% of enterprise generative AI pilots delivering no measurable financial return.
The reason that matters to you specifically: Dataiku and the Harris Poll, surveying 900 CEOs across eight countries, found 80% of CEOs saying their job is at risk if AI fails, and 87% saying they would stake their job on delivering results from their AI initiatives. The pressure you are feeling is not unusual and it is not a signal that you handled this badly.
What is unusual is running a proper post-mortem instead of quietly starting another pilot. Almost nobody does it, which is why the second pilot usually fails for the same reason as the first.
The seven-question post-mortem
Work through these in order and stop at the first one where the honest answer is uncomfortable. That is your cause.
- Why this process, and who chose it?If the answer traces back to a demonstration somebody saw, you have found it. The justification was assembled after the fact to fit a solution that already existed. This is the single most common cause.
- What did one occurrence of that process actually cost, before we started?If nobody can answer, the project never had a business case, only a hypothesis. You cannot measure a return against a baseline you never took.
- Did the same field mean the same thing in every system it came from?Not whether the data existed. Whether it was consistent. Inconsistent definitions are learned faithfully by any model, and this is discovered late because everyone assumes their own data is fine.
- Who was going to accept the risk of putting it in front of customers?Name the person. If you cannot, the pilot was never going to ship regardless of how well it performed, and it became a permanent pilot by default.
- What number did we agree in advance would count as working?Agreed before the build, in writing. Without it, evaluation becomes an argument about impressions that nobody can win, and the project dies of ambiguity.
- Did anyone whose work it changed have a say before it was built?Systems that arrive as a surprise get worked around. This is not resistance to change, it is a design failure that was avoidable.
- Was this the first thing we ever put into production?If so, you attempted the hard part of the use case and the hard part of operating AI simultaneously. That combination produces most pilot graveyards.
Which failure is yours?
| What you observed | Actual cause | What has to change |
|---|---|---|
| It worked and nobody used it | Wrong process chosen | Re-select on cost and frequency, not on how well it demos |
| Good in testing, poor in reality | Inconsistent data definitions | Fix definitions before touching a model |
| Endless further evaluation rounds | Nobody would accept go-live risk | Name an accountable owner and define the rollback |
| Legal or compliance stopped it | Governance considered too late | Design oversight in from the start |
| Staff quietly worked around it | Nobody affected was consulted | Involve them before the build, not at training |
| Ran out of budget mid-build | Scope was never bounded | Fixed scope with a defined end date |
| Vendor delivered what was asked, it was wrong | Specification error, not delivery error | The specification is the work |
What to tell the board
The instinct is to minimise. That is the wrong move, because a second failure after a minimised first one is much harder to survive than an honestly reported first one.
- Lead with the cause, not the outcome. "We chose the wrong process and here is how we now choose" is a management answer. "It did not work out" is not.
- Give the industry context once, briefly, and move on. More than 80% fail. Use it to establish that this is a known pattern, not to excuse anything.
- Say what it cost and what was learned. Both numbers. The learning has to be specific enough to be checkable.
- Bring the revised selection method. Not a new project. The method by which the next candidate will be chosen.
- State the stopping rule. At what point you would stop the next one, decided in advance. Boards trust plans that can be abandoned far more than plans that cannot.
Should you try again, or stop?
A genuine answer, including the case for stopping.
- Try again if you can now name a process that is frequent, expensive and unloved, cost one occurrence fully loaded, and multiply it out to an annual number that is obviously larger than the whole cost of the work.
- Try again if the failure was ownership or data definition, both of which are fixable and cheaper to fix the second time.
- Stop for now if no process survives that arithmetic. There is no shame in it and it is more common than the market admits.
- Stop for now if the underlying process is broken. Automation will make a broken process faster and no better, and you will have paid to accelerate the problem.
Deciding to stop is a legitimate outcome of a post-mortem. It is also the outcome nobody selling you an AI project will ever recommend.
Frequently asked questions
How common is it for an AI pilot to fail?
Very. RAND reports that by some estimates more than 80% of AI projects fail, about twice the rate of IT projects that do not involve AI, and MIT NANDA found 95% of enterprise generative AI pilots delivering no measurable financial return. A failed pilot is the normal outcome rather than the exception.
Why do AI pilots fail?
Rarely because of the model. The recurring causes are choosing a process because it demonstrated well rather than because it was expensive, discovering late that data definitions differ across systems, and having nobody willing to accept the risk of putting the system in front of customers.
Should we try a different AI model or vendor?
Usually not. A different model solves a capability problem, and most failures are selection, data or ownership problems. If the process was not worth automating, a better model produces a better answer to a question nobody was asking.
What should we tell the board after a failed AI pilot?
Lead with the cause rather than the outcome, state what it cost and what was learned, present the revised method for selecting the next candidate rather than a new project, and define in advance the point at which you would stop.
How do we choose a better second project?
Rank candidates by annual frequency multiplied by fully loaded cost per occurrence, discount by the share genuinely automatable, and compare against the entire cost of specification, build and running. The right second project is usually boring and high-volume.
Is it reasonable to stop trying with AI?
Yes, and it is an underrated outcome. If no process survives the cost arithmetic, or if the underlying process is broken, stopping is correct. Automation applied to a broken process makes it faster and no better.
If you are writing the post-mortem now
The seven questions above are the whole method. If you want a second opinion on what they surfaced, that is a short conversation.