False positive or missed answer: scope which error you fear

A fraud lead at a payments company kept asking his team to make the model more accurate. The question that actually mattered was which kind of mistake he could live with, and he had never named it.

False positive or missed answer: scope which error you fear

A fraud lead at a payments company told me his team had been chasing accuracy for months and getting nowhere useful. Every review meeting was the same. The model was 92 percent accurate, then 93, then a config change dropped it back to 91. The number moved and nothing felt better. He was frustrated because the metric was improving and the outcomes were not.

The problem was that “accuracy” was hiding the only question that mattered. His model made two completely different kinds of mistakes, and they had nothing in common except that both counted against the accuracy number. Sometimes it flagged a legitimate transaction as fraud. That was a false positive, and it meant a real customer got their card declined at a checkout, called in angry, and sometimes left. Sometimes it missed real fraud. That was a missed case, and it meant money actually walked out the door.

Those two errors do not cost the same. They do not even cost the same kind of thing. One costs customer trust and support load. The other costs money directly. A single accuracy number treats them as interchangeable, and they are not. Until he decided which one he feared more, no amount of tuning could point in a direction.

One number hides the trade-off

Almost every AI decision system faces this. There is a knob, sometimes literal, sometimes buried in the design, that trades one error against the other. Turn it one way and you catch more fraud but also decline more good customers. Turn it the other way and you stop annoying good customers but let more fraud through. You cannot minimize both at once. The knob only slides.

That means the real scoping question is not “how accurate can we make it.” It is “which mistake are we willing to make more of, so we make less of the one that hurts more.” That is a business decision, not a modeling decision, and it has to be answered before anyone tunes anything.

For his fraud case, the honest answer took a hard conversation. Missing fraud cost money that was easy to count. Declining good customers cost trust that was harder to count but, they eventually agreed, worse over time. So they scoped the system to lean toward catching fraud, accepting more false flags, and paired it with a fast human review path so wrongly declined customers got unblocked quickly. The design followed the fear, once the fear was named.

The costs are rarely symmetric

The two errors almost never cost the same, and the whole point of discovery is to find out which way the scale tips.

  • In fraud, a missed case loses money and a false positive loses a customer. Usually you fear the missed case, but not always.
  • In medical screening, a missed case can be a missed diagnosis and a false positive is an unnecessary follow-up test. There you almost always fear the missed case.
  • In content moderation, a false positive removes something legitimate and a missed case leaves something harmful up. Which you fear depends entirely on the platform and its stakes.

A head of risk at a lending startup put it plainly to me. His whole model design came down to one sentence he made the team write on the wall: we would rather decline ten good applicants than approve one that defaults. Once that sentence existed, every threshold decision had an answer. Before it existed, every threshold decision was an argument.

Name the error you fear, and the tuning stops being a guessing game.

How we approach it at Density Labs

In the AI Opportunity Assessment, our two-week fixed engagement at $2,500, we make teams say out loud which mistake costs more before we look at any model. A false positive and a missed case get separate names, separate owners, and separate costs, because a single accuracy target quietly averages away the exact decision that should drive the design. Getting that sentence written down is often the most valuable hour of the engagement, and it costs a conversation, not a quarter.

Stop asking how accurate the model is. Ask which mistake you would rather make, and design toward that answer.