The stale document in your knowledge base is a confident wrong answer waiting

One outdated file that should have been retired years ago is still in the index. The day retrieval picks it, the model grounds a wrong answer in it and sounds completely sure.

The stale document in your knowledge base is a confident wrong answer waiting

Every knowledge base has one. A document that was replaced, contradicted, or simply made wrong by a decision last year, and that nobody ever pulled from the index. It sits there quietly until the day retrieval picks it, and then it becomes the source of an answer delivered with total confidence.

Adding is easy, retiring is nobody’s job

A product and operations leader who now consults on AI for mid-sized companies starts every engagement from the business problem, not the technology, and is willing to tell a client the AI is not the fix. That discipline points at something teams skip. They build the knowledge base by adding. New document, into the index. Revised policy, into the index. What almost never happens is the opposite move. Old document, out of the index. Retiring content is a decision someone has to make and own, and it has no deadline, so it loses to everything that does.

So the corpus only grows. The current version of a policy goes in. The previous version stays. Now the index holds two documents that both talk about the same subject with equal authority, and only one of them is true. Retrieval cannot tell which is which from meaning alone, because the stale one is often the better semantic match. It was written carefully. It uses all the right words. It just describes a world that no longer exists.

The wrong answer that outranks the right one

Here is the scenario that costs someone. A rule changed. The old rule allowed something the new rule forbids. Both documents are in the index. A user asks about that exact rule. The old document, being detailed and on-topic, gets retrieved. The model reads it and answers with the old rule, in clear and grounded prose, citing a real internal source. The user acts on it. The answer was authoritative, traceable, and wrong, and every signal the system produced said it was fine.

That is the danger of a stale document. It does not degrade the output into something obviously broken. It produces a clean, confident, well-sourced answer that happens to be from the wrong version of reality. There is a rough reliability bar that AI features need to clear, something like 85 percent accuracy before they can be trusted, and a corpus full of un-retired documents chips at that number from a direction no model tuning can reach. And the broader picture is not encouraging. Studies of enterprise GenAI put the share of pilots that deliver no measurable return around 95 percent, and grounding on outdated content is one of the quieter ways a promising pilot ends up in that group.

Content needs an expiry, not just a birthday

The teams that handle this give documents a lifecycle. A document does not just get created and indexed. It gets a review date, an owner, and a way to be marked superseded so retrieval stops considering it. When a policy is replaced, the old one is expired in the same motion, not left as a landmine. This is unglamorous curation work, and it is exactly the work that decides whether your knowledge base is a source of truth or a pile of every version of the truth you ever had.

How we approach it at Density Labs

In the AI Opportunity Assessment, our fixed two week, $2,500 engagement, we look for the documents that should have been retired and were not. We check whether your content has a way to expire, whether superseded versions can still be retrieved, and where two documents in the index disagree with each other. Finding the stale file before it grounds a wrong answer costs a fraction of what it costs to explain, after the fact, why the AI confidently cited a policy that ended last spring.

An index that never forgets is an index that cannot tell you what is still true.