Success criteria have to be falsifiable or the pilot never ends
A head of growth at a B2B SaaS company ran an AI pilot with no way to fail. It could not end, because there was no line it could fall below. She started writing criteria that could lose.
Success criteria have to be falsifiable or the pilot never ends
A head of growth at a B2B SaaS company told me about an AI pilot that ran four months longer than it should have because it could not fail. The goal, written in the plan, was to “improve the onboarding experience.” The pilot launched. It did something. And then nobody could say whether it had worked, because there was no line it could fall below. Every review became an argument. Supporters pointed at the parts that looked good. Skeptics pointed at the parts that did not. The decision to keep going or stop had nothing to lean on, so it leaned on whoever was most senior in the room that week.
She learned the hard way that a success criterion you cannot fail is not a criterion. It is a mood. And a pilot measured by mood never ends, because you can always find a reason to run it one more month.
If it cannot lose, it cannot win
Her rule now is that a success criterion has to be able to come out false. “Improve the experience” cannot be false. Whatever happens, someone will say the experience improved somewhere. That is exactly why it is useless. A real criterion names a thing you can measure, a direction, and a line, decided before the pilot starts, so that when the number lands, the answer is already agreed.
She gave me the before and after. Before: improve onboarding. After: new users who see the AI setup guide reach their first configured project at a higher rate than users who do not, within their first session, and if they do not, we stop. That version can fail. If the number comes back flat, the pilot is over, and nobody has to win an argument to end it. The criterion ended it.
The falsifiable version does something the vague one cannot. It forces the disagreement to happen before the work, when it is cheap, instead of after, when it is loaded with sunk cost and someone’s reputation. If two people cannot agree on the criterion up front, they were never going to agree on the result. Better to find that out in a planning meeting than in a launch review.
Write the line that ends it
Her checklist before any AI pilot is short and a little uncomfortable, because it asks people to name the number that would kill their idea.
- A metric that is already measured, or can be, without a special project to define it.
- A direction and a line, so the result is a yes or a no, not a discussion.
- A comparison, what this is being measured against, so “better” means something.
- A pre-agreed action for each outcome, including the one where you stop.
The last item is the one people skip and the one that matters most. If you have not agreed in advance what a failing number means, a failing number will just start a negotiation. Write down “if we do not clear this line, we stop” before you launch, and the number gets to decide instead of the org chart.
A friend who runs product analytics at a marketplace told me she will not let a pilot start until someone can describe the result that would make them shut it down. If no one can, she says, the pilot is not a test. It is a purchase everyone has already decided to keep, and they should stop pretending it is a test. AI features usually need to clear about 85 percent accuracy before their mistakes stop compounding, but even a clean accuracy number means nothing if you never agreed what business result it had to produce.
How we approach it at Density Labs
In the AI Opportunity Assessment, our two-week fixed engagement at $2,500, we make the team write a success criterion the pilot can fail, before we scope the build. We push for a metric that already exists, a line agreed by the person who owns the outcome, and a stated action for the failing case. It is a short conversation and occasionally a tense one, because it surfaces disagreement early. That is the point. Better to have the fight over the criterion than over the result.
Write a success criterion your pilot can actually fail, and agree it before you start. If it cannot lose, it will never end.