# When to kill an AI pilot: the numbers that should stop you

_A pilot that will not die is more expensive than one that fails fast. The discipline is deciding the kill numbers before you are emotionally invested in the answer._

The numbers that should stop an AI pilot, and why you have to set them before you are attached to the outcome. An anonymized story on early-warning signals and reading failure before it is catastrophic.

# When to kill an AI pilot: the numbers that should stop you

Killing an AI pilot is harder than starting one, because by the time the evidence is in, people are attached. The engineer believes in it. The sponsor has told their boss about it. So the pilot limps along past the point where the numbers already said stop. The fix is to decide the kill numbers early, while you can still be objective, and then honor them.

## Read the early signal, not the catastrophe

The founder of a predictive-maintenance company gave me a useful mental model. His whole business is catching failure in critical equipment months ahead of time, because, as he put it, infrastructure works until it does not. Most operators wait for the catastrophic event. His sensors read the early anomaly, the quiet signal that a failure is coming, so someone can act while intervention is still cheap. He also spends real effort separating true faults from noise, because a system that cries wolf gets ignored.

An AI pilot has the same early signals if you decide to read them. The catastrophic version is a pilot that consumes two quarters and a budget before anyone admits it failed. The early-signal version reads the leading indicators in weeks and calls it. The skill is the same as his: know which numbers are a real fault and which are noise, and act on the real ones early.

## The numbers that should stop you

Set these before the pilot, when you are calm. A quality floor: if the output cannot clear the level the task actually requires after honest iteration, and it is not trending toward it, that is a stop. A cost ceiling: if cost per task at production volume breaks the economics even when quality is fine, that is a stop, and it is one of the most common ones, because a pilot measures feasibility, not ROI, and infeasible unit economics do not improve at scale. A feasibility wall: if the feature cannot meet the latency or integration reality it has to live in, no amount of model tuning saves it.

And a null baseline. If you set the baseline before deployment and the pilot cannot show movement against it, the ROI claim is unfalsifiable, and an unfalsifiable pilot that has run long enough to prove itself and has not is a stop by default.

## Separate a kill from a pause

Not every red number means dead. Some mean pause and fix a specific input. The way to tell them apart is to have written, before the pilot, which numbers are fatal and which are fixable. Use a holdout group or a controlled before-and-after so the signal you are acting on is real and not seasonal noise, the same discipline that keeps a maintenance alert from being a false alarm. Deciding the kill conditions in advance is what lets you shut down a pilot without it becoming a referendum on the person who built it.

## How we approach it at Density Labs

Our AI Readiness Assessment is a fixed two week engagement priced at $2,500. We write the kill conditions next to the success conditions, before the pilot, in numbers: the quality floor, the cost ceiling, the feasibility wall, and the point at which no movement against baseline means stop. We define how you will tell a fatal signal from a fixable one, so the decision at the end is mechanical and not emotional.

A pilot that cannot fail cannot teach you anything. Decide the stop numbers first, and a dead pilot dies cheap.
