Why your pilot's 95% felt fine and production's 95% did not
The accuracy number held steady from pilot to launch. The experience fell apart anyway. The number was never the thing that changed.
Why your pilot’s 95% felt fine and production’s 95% did not
The pilot hit ninety-five percent accuracy and everyone was happy. The feature shipped. The dashboard still says ninety-five percent. And the complaints are pouring in. The number held. Something else did not, and the number was never going to tell you what.
The same figure, a different world around it
Ninety-five percent in a pilot and ninety-five percent in production are the same fraction over two very different denominators, measured by two very different audiences.
A director of engineering at a logistics company walked me through this after a rough launch. In the pilot, the AI classified inbound shipping documents, and the five percent it got wrong were reviewed by the two analysts running the trial. They corrected the misses quietly, learned the model’s blind spots, and reported a number that felt great. Then it went live across every branch. The five percent was now spread across people who had never seen the tool, had no idea it could be wrong, and treated its output as fact. Same accuracy. The trust around it had changed completely, and misplaced trust turns a known error rate into a series of surprises.
MIT’s 2025 State of AI in Business study found that roughly ninety-five percent of enterprise GenAI pilots deliver no measurable return. A lot of those pilots posted fine accuracy numbers. The number was rarely the reason they stalled.
The errors move when the inputs move
There is a second reason the same percentage feels worse in production. The mix of what the model sees changes.
A data lead at a healthcare company described it precisely. During the pilot, the inputs came from a few cooperative departments with clean, consistent records. The model looked accurate because the data was easy. Production opened the feature to the whole organization, and the incoming records were messier, older, and formatted a dozen ways nobody had standardized. The accuracy figure barely moved, because the easy cases still dominated the count. But the hard cases were now real, they were new, and they were exactly the ones users noticed. The average stayed calm while the felt experience got worse.
She made a rule out of it. Never trust a pilot accuracy number that was measured on data the pilot got to choose.
Where the gap actually lives
When a steady accuracy number produces an unsteady launch, look here first:
- Who is watching. Pilot errors get caught by invested people. Production errors reach users who assume the output is right.
- What data flowed in. A cooperative pilot sees clean inputs. Production sees everything, including the records nobody cleaned.
- What one error means now. A wrong answer to a tester is a note. A wrong answer to a customer is a broken promise.
- What you measured. A single top-line percentage averages away the segments where the failures actually cluster.
How we approach it at Density Labs
When we help a team read a pilot result, we spend as much time on the conditions the number was measured under as on the number itself. Who reviewed the outputs. Which data the pilot was allowed to see. What happens to a wrong answer once no one is invested in catching it. The accuracy figure is real. It just describes the pilot’s world, and production is a different world wearing the same percentage.
Ninety-five percent was never a property of your model alone. It was a property of your model, your data, and your reviewers together. Change the last two and the first one lies to you.