The number gets quoted so often it has stopped meaning anything. The useful part is not that most pilots fail. It is where they fail, because the pattern is remarkably consistent and almost none of it is about the model.
In the post mortems we have run and sat in on, dead pilots cluster on the organizational side by roughly nine to one. The demo worked. The model was fine. What was missing was an owner, data anyone could actually use, and a plan for the last mile into the systems people work in every day.
| Failure | What it looks like | The tell, before you start |
|---|---|---|
| No owner | A successful demo with no one whose job depends on it shipping. It survives until that person's quarter gets busy. | Nobody can name the single person accountable for the outcome, only for the project. |
| Data not ready | The model works on the extract someone hand cleaned. Production data is inconsistent, permissioned differently, or arrives late. | The pilot dataset was prepared by hand, once, by a person who is not on the build team. |
| The last mile | Output lands in a dashboard nobody opens instead of the CRM, ticket, or queue where the work actually happens. | No one has said which existing screen changes. |
| No decision criteria | The pilot cannot fail, so it cannot succeed either. It just continues. | There is no number written down that would end it. |
Frontier models are now good enough for the overwhelming majority of mid market workflows. That is a real shift from three years ago, and it moves the bottleneck. When the model is a commodity, the differentiator is everything around it: who owns it, what it reads, where it writes, and whether anyone changed their behavior because of it.
This is why swapping models rarely rescues a stalled pilot. If a pilot died of no ownership, a better model produces a better artifact that still nobody owns.
| Pilots that stall | Pilots that ship | |
|---|---|---|
| Scope | Explore what AI could do for us | One workflow, named, with a number attached |
| Ownership | A project sponsor | An operator whose weekly work changes |
| Data | Cleaned by hand for the demo | Read from the real source on day one, warts included |
| Integration | A new dashboard | An existing screen people already open |
| Exit | Open ended | A written go or no-go date and threshold |
| Team | Handed to whoever has capacity | A senior engineer who will also run it in production |
Every failure above is visible in advance. That is the entire argument for scoping properly before building: the questions that predict failure are cheap to ask and expensive to skip. A two week diagnostic answers them for less than the cost of a month of a stalled pilot.
Sometimes the honest answer is do not build this. That is a good outcome. The expensive version is finding out in month five.
Related reading: what an AI diagnostic actually includes · running the pilot in house vs bringing in a team · the AI Readiness Assessment