Business & Money
Why Do Most AI Pilots Stall, and How Do You Run One That Ships?
By Jim Vernon, Editor, AI Intelligence International · Published 28 August 2026 · Reviewed against our editorial standards · About the author
The common pattern is a pilot that demos well, generates enthusiasm, and then quietly stops. Nothing failed technically; the pilot simply never had an owner, a budget line or a defined production standard on the other side of it.
This article covers the structural decisions that determine whether a pilot ships, most of which have to be made before the pilot starts rather than after it succeeds.
Key takeaways
- A pilot without a named production owner is a demo, regardless of results.
- Define the pass mark before you see any output, or the goalposts will move.
- Pick a process with an existing baseline; you cannot show improvement against an unmeasured process.
- Budget the production run before the pilot, because that is the approval that actually blocks.
What makes a pilot die after a successful demo?
The handover gap. Pilots are typically run by an enthusiastic team with slack capacity, and production requires a team with a budget, an on-call rotation and an owner who will be accountable when it misbehaves.
If nobody agreed in advance to be that owner, the pilot result lands on a leadership team who now have to find one, and the path of least resistance is to commission another pilot.
Name the production owner at kickoff and have them attend the pilot reviews. Their questions will be different, and better, than the pilot team's.
How do you choose which process to pilot?
Pick one that is already measured. If you cannot state today's error rate, cycle time or cost per unit, you will not be able to demonstrate improvement, and the debate will collapse into opinion.
Prefer high volume and low individual stakes for the first pilot. Volume gives you a statistically meaningful result in weeks rather than quarters; low stakes keeps the risk of a bad output survivable.
Avoid processes with heavy regulatory exposure for a first attempt, not because they are unsuitable but because the approval overhead will dominate the timeline and teach you nothing about the technology.
What should the pass mark be?
Written down before any output is seen, with a number and a comparison. 'Matches current quality on 90% of cases at under 40% of current cost per case' is a pass mark. 'Shows promise' is not.
Include a floor as well as a target. Define the result that would make you stop, and commit to stopping. Pilots without a stop condition rarely end; they fade.
Get the pass mark signed by the production owner and the budget holder, not just the project team. That is what prevents renegotiation once results arrive.
How long should a pilot run?
Long enough to see the tail. Four to six weeks is typical for a high-volume process; a two-week pilot mostly measures the easy cases because unusual inputs cluster over longer windows.
Set the end date at the start and do not extend it. Extensions are almost always a sign that the pass mark was not met and someone is hoping for a different result.
If the process is seasonal, either wait for a representative period or explicitly document that the result does not cover peak conditions.
What has to be true for production, and who pays?
Production needs monitoring, an escalation path, a rollback, an owner, and a recurring cost line. Every one of those costs money that a pilot does not.
Estimate the annual run cost at kickoff — licences, infrastructure, and the human review time that almost never goes to zero — and get provisional approval for it conditional on the pass mark.
This is the single highest-leverage step. A pilot that clears its pass mark against a pre-approved budget ships in weeks; the same pilot without one enters a funding cycle and loses momentum.
How do you measure honestly during the pilot?
Sample the outputs rather than reviewing all of them, and have the sampling done by someone who did not build the pilot. Self-review by the pilot team reliably overstates quality.
Track corrections by category, not just a pass rate. Knowing that 8% of outputs failed is less useful than knowing that 6 of those 8 points came from one input type you could route away.
Record the human time spent per case including review. Pilots frequently report a large saving in generation time and quietly ignore the review time that replaced it.
Worked example: two pilots at the same company
A logistics firm ran two AI pilots in the same quarter. The first targeted freight claim summarisation, chosen because it was interesting. The second targeted delivery exception categorisation, chosen because it already had a weekly volume figure and a known handling time of 6.5 minutes per exception.
The claims pilot had no pass mark and no production owner. It demoed well in week three, generated a well-received deck, and had no budget line. It was still 'awaiting prioritisation' two quarters later.
The exceptions pilot set a pass mark of matching human categorisation on 92% of a blind sample at under half the handling time, with the operations manager as named production owner and a provisional run budget of 24,000 per year.
It ran five weeks over roughly 4,100 exceptions. Blind review by a supervisor put agreement at 94%, with almost all disagreements in one carrier's malformed messages, which were routed to humans. Effective handling time fell to 2.4 minutes including review.
It shipped eleven days after the pilot ended, because the only remaining decision was whether the pass mark had been met.
Frequently asked questions
How much should a first pilot cost?
Keep it small enough not to need a formal business case — often a few thousand in tooling plus a few weeks of part-time effort. The expensive approval is the production run, so spend your political capital there rather than on the pilot.
Should the pilot team build the production version?
Not usually. Pilot code optimises for learning speed and production needs monitoring, error handling and maintainability. Plan for a rebuild and budget it, rather than being surprised by it.
What if the pass mark is nearly met?
Treat near-misses as failures for the stated scope, then consider a narrower scope where the mark is comfortably met. Shipping a smaller thing that works beats shipping a marginal thing everyone distrusts.
How many pilots should run at once?
For most organisations under a few hundred people, one. Parallel pilots split the attention of the small number of people who can actually evaluate results, and produce two half-answers instead of one decision.