How long should an AI pilot run before you kill it or scale it?
Too short and you kill something that was about to work. Too long and you fund a demo forever. The answer is a window with a decision at the end of it.
How long should an AI pilot run before you kill it or scale it?
The wrong way to run an AI pilot is to leave the end date open. With no window and no decision point, one of two things happens. You kill it early on a bad week, or you fund it forever because nobody is allowed to call it. Both outcomes are failures, and both are avoidable with one rule: a pilot is a window with a decision at the end of it.
The reason this matters is the base rate. MIT’s 2025 study found about 95% of enterprise GenAI pilots delivered no measurable return. A pilot with no decision point does not escape that number, it just takes longer and costs more to join it.
Give the signal a real window
A consultant who builds diagnostic assessment tools for expert businesses gave me a clean example of a bounded window producing a real answer. She turned one client’s modest assessment, priced at $14.99, into $70,000 in its first 60 days, organically. What makes that useful here is not the revenue. It is the shape. A small, cheap thing was given a defined window, and inside that window it produced a signal clear enough to act on. Scale it, obviously.
That is the model for an AI pilot. Keep it small and cheap, but give it a real window and a metric it has to move inside that window. Sixty days is not magic, but the structure is. A defined period, a number that decides the outcome, and a genuine willingness to act on what you see, up or down.
She also works from behavioral data rather than opinion, letting what people actually do guide the next move. A pilot deserves the same honesty. The decision at the end should rest on what the metric did, not on how attached anyone got to the demo. The most expensive pilots are the ones kept alive by enthusiasm long after the signal said stop.
The window length itself should match the metric you chose. If the outcome you care about is something a user does in a session, a 30-day window may be plenty to read the signal. If it is a slower behavior that only shows up over repeated use, 30 days will lie to you and you need 60 or 90. The mistake is picking a window out of habit and then reading a signal the window was never long enough to produce. Decide what you are measuring first, and let that set how long you watch.
Setting the window before you start
A pilot with a clean kill-or-scale decision has these defined on day one.
- A fixed window. A specific length, often 30 to 90 days, not “until it feels ready.”
- A decision metric. The one number that has to move for this to be worth scaling, agreed before you begin.
- A baseline. What the metric is now, so the change is real and not a story you tell afterward.
- A default answer. If the metric does not move, the pilot stops. Silence is not a reason to continue.
Without these, a pilot has no end, and a pilot with no end is a budget leak wearing a demo.
How we approach it at Density Labs
When a team is unsure whether to keep going or pull the plug, we use the AI Readiness Assessment, our $2,500 engagement, to set the window and the decision before more money goes in. We define the metric, capture the baseline, and agree the length, so the kill-or-scale call is made on evidence instead of attachment.
That diagnostics consultant proved a bounded window can surface a decisive signal in 60 days. An AI pilot can do the same, but only if you decide up front what the window is for.
Give it a real window, then be honest about what the window showed.