# What to measure in the first 30 days of an AI pilot

_Month one is too early for ROI and too important to waste. Measure the signals that tell you whether this thing can survive contact with production, not the ones that flatter it._

The right things to measure in the first 30 days of an AI pilot: feasibility signals, not return. An anonymized solo-builder story on testing everything in a closed alpha, and the leading indicators to track.

# What to measure in the first 30 days of an AI pilot

The first 30 days of an AI pilot are where most of the real information is, and where most teams measure the wrong thing. They go looking for ROI, which is not there yet, and miss the feasibility signals, which are. Month one is a feasibility test. Measure whether the thing can stand up, not whether it made money.

## A closed alpha is a measurement instrument

An engineer building an AI product solo described his early phase in a way that captures the point. He runs an invite-only alpha where, in his words, he is testing out everything. The small closed group is not a soft launch, it is an instrument. He watches how the system behaves under real use, where it breaks, whether it holds up when several people lean on it at once, and how it performs under load before he lets anyone else near it.

That is exactly the posture for the first 30 days of a pilot. You are not trying to impress anyone. You are trying to find out what falls over. A closed group, heavy real usage, and honest attention to the failures is how you learn whether the pilot has a future, long before it could possibly show a return.

## Measure feasibility, not return

The frame that keeps month one honest: a pilot measures feasibility, not ROI, and expecting a return number this early guarantees a false negative. In the first 30 days, track the signals that predict whether this can reach production. Quality on real inputs, not the curated demo set. Latency under realistic load. Cost per task at actual volume, extrapolated forward. Failure and escalation rate. How often a human has to step in and redo the output. These are leading indicators. They tell you where the pilot is heading while there is still time to steer.

Do capture the business baseline in this window too, because it is the only clean moment to get it. Set the baseline before deployment or the eventual ROI claim is unfalsifiable, and the first 30 days is when "before" still exists. Every day the pilot runs without that baseline captured is a day of evidence you cannot get back, so treat baseline capture as a week-one task, not a later one.

## Watch the numbers that would kill it

Month one is also when you find the disqualifiers. A cost per task that only works at pilot volume and explodes at production volume. A quality level that looks fine on average but fails badly on the inputs that matter most. A latency budget the workflow cannot absorb. These are cheaper to discover in week three than in month six, and the whole reason to measure them early is to give yourself permission to stop before the sunk cost gets heavy. A feasibility test that surfaces a fatal number in 30 days did its job.

## How we approach it at Density Labs

Our AI Readiness Assessment is a fixed two week engagement priced at $2,500. We define the first-30-days measurement plan: the feasibility signals to track, the load and cost tests to run against realistic volume, the failure and escalation rates to watch, and the baseline to capture while it is still available. We separate the leading indicators from the return so nobody judges a month-old pilot on a number it cannot have yet, and so a real red flag gets caught while it is still cheap.

Month one is a stress test, not a sales pitch. Measure what breaks, and the pilot tells you early whether it is worth the rest of the year.
