The pilot scoped to a slice that does not represent production
A data lead at an agri-tech company ran a clean, convincing AI pilot. Production data looked nothing like the pilot slice, and the model that shone in the test struggled the moment it went live.
The pilot scoped to a slice that does not represent production
A data lead at an agri-tech company told me about a pilot that went beautifully and then betrayed him. The model predicted crop issues from field sensor data, and in the pilot it was sharp. Good accuracy, clear signal, the kind of result that gets a project funded. The pilot ran through one growing season on a set of well-instrumented farms, and everyone walked away confident. Then it went to production and the accuracy fell apart in ways nobody had seen coming.
The model was fine. The pilot was the problem. It had been scoped, without anyone deciding to, on a slice of data that looked nothing like what production would actually feed it.
The slice was clean, seasonal, and narrow
Three things made the pilot data unrepresentative, and each was invisible until launch. The pilot farms were the well-run ones with maintained sensors, so the inputs were clean. Production included farms with patchy, drifting, sometimes broken sensors, and the model had never seen noise like that. The pilot ran through a single season, so the model learned one set of weather and growth patterns. Production hit conditions the pilot never contained. And the pilot farms were a narrow band of crop types and soil, while production spanned a much wider range. The pilot had quietly handed the model an easy version of the problem and called the score a prediction of the real one.
This is the trap with pilots, and it is common because the easy slice is the available slice. When you set up a pilot, you naturally reach for the data that is clean, recent, and close at hand. That data is easy precisely because it is not representative. The messy, varied, long-tail data that production will actually deliver is harder to assemble, so it gets left out. Then the pilot measures performance on the easy version and everyone reads it as performance on the real one.
A friend who leads analytics at a retail company described the identical surprise. Her team piloted a demand-forecasting model on a few flagship stores, which were large, stable, and well-staffed. It worked wonderfully. Rolled out to the small and irregular stores, which were most of the chain, it stumbled, because those stores behaved nothing like the flagships the pilot had chosen. The pilot had answered a question nobody was actually asking.
Scope the pilot to look like production
The correction is to treat representativeness as a design requirement of the pilot, not an afterthought. Before you run it, ask how production data will differ from what you are about to test on, and deliberately pull the hard cases into the pilot. You want the noise, the seasons, the range, and the long tail in the test, even though including them makes the pilot look worse. A pilot that looks worse but resembles production is telling you the truth. A pilot that looks great on a clean slice is telling you a story.
Questions worth asking before you trust a pilot result:
- How will production inputs differ from the pilot data in quality, range, and time?
- Are the pilot cases the easy ones, and if so, why were they available?
- Does the pilot span the full variety production will hit, or a narrow band of it?
- What conditions will production see that the pilot window never contained?
If the pilot data is cleaner, narrower, or more seasonal than production, the encouraging number is not a forecast. It is a best case you will not get to keep.
How we approach it at Density Labs
Checking a pilot for representativeness is part of the AI Opportunity Assessment, our fixed two-week engagement at $2,500. Before anyone reads a pilot result as a green light, we map how production data will differ from the test data and make sure the hard cases are in the pilot, not hiding outside it. Sometimes the finding is that a promising pilot proved almost nothing because it ran on the easy slice, and that reframe saves a client from launching on a number that was never going to hold.
A pilot on a clean, narrow, seasonal slice tells you how the model does on the easy version of the problem. Scope the pilot to look like production, because production is the only version that ships.