Back to blog
AI at Work

What Makes an AI Answer Trustworthy at Work

5 min read
AI knowledge baseinternal AI knowledge baseAI that answers questionscompany knowledge AItrustworthy AI answers

The short version

  • An answer is trustworthy at work when you can see where it came from and who owns it. Everything else is a guess with good grammar.
  • Accuracy rates are the wrong measure. What matters is whether the system tells you when it does not know.
  • Four properties to require: a source under each answer, a named owner, a date, and an explicit "not in the record" response.
  • Test any tool with questions you know are unanswerable. That is where the difference shows.

Teams adopting an internal AI knowledge base run into the same wall at about week six. The answers are good. Nobody can tell which ones to rely on. So people either check everything, which removes the benefit, or they stop checking, which is worse.

Why a confident answer is a problem

A colleague who does not know something says so. They hedge, they say they think it was around March, they suggest asking someone else. That hedging is information, and you use it without noticing.

A system that produces fluent prose regardless strips out that signal. The answer built from three solid sources and the answer assembled from one loosely related page look identical. You cannot tell them apart, so you have to treat all of them the same, and the only safe way to treat all of them the same is with suspicion.

That is why adoption plateaus. It is not that the tool is wrong often. It is that you cannot tell which times.

Four properties of a usable answer

Property Why it matters
A source under each answer Lets you verify in seconds rather than redo the search. Without it, verification costs as much as the original question.
A named owner Tells you whose understanding this is, so you know whom to go to when it is wrong or out of date.
A date Most workplace answers have a shelf life. An answer from March presented without its date is a trap.
An explicit "not in the record" The property everything else depends on. Without it, the other three decorate answers that should not have been given.

The fourth is the hard one, technically and commercially. A system that declines to answer looks less capable in a demonstration, and it is the only kind you can put in front of a client conversation.

Why accuracy is the wrong measure

A system that answers everything with ninety percent accuracy sounds better than one that answers sixty percent of questions and says so for the rest. In practice the second is far more useful.

With the first, every answer carries a hidden probability of being wrong, and you have no way of knowing which. So the value of all of them is capped by your willingness to act on the worst.

With the second, the answers you get are reliable and the non answers route you to a person. You can act on what you receive. The coverage is lower and the usable value is higher, which is not how tools in this category are usually sold. See when an AI should decline to answer.

What the answers should come from

The source material determines the ceiling. Most internal AI knowledge bases are built over documents, which means they inherit the documents' problem: what people ask about is current state, and current state is rarely written down. See most of that searching is looking for a person.

This is what StandIn is built on, and the design follows the four properties above. Each day a person spends about ninety seconds on a brief: what moved, what is open, what is blocked, what is next, mostly drafted from the work that already happened. They confirm it, which is what gives every later answer a named owner and a date. When they are unavailable, their StandIn answers from that brief, in their words, with a source under every answer. It never guesses. If the answer is not in what they wrote, it says so and names who to ask.

The trade is deliberate. It knows less than a system indexing everything your company has ever produced, and what it knows is current, attributable and checkable. For questions you would otherwise have asked a colleague, that is the property that matters. See what it will not do and grounding AI in company decisions.

How to test it

Write twenty questions before any trial, in three groups.

  • Ten you know are answerable from what your company has written. This checks basic competence.
  • Five where the answer changed recently. A system that returns the old answer without a date is the most dangerous failure mode, because it is confidently wrong about something that was once right.
  • Five you know are not written down anywhere. This is the group that decides it. Anything other than an explicit "I do not have that" is a fail.

Run the same twenty against every option and score only the last ten. The first ten will look similar across vendors.

Common Questions

Does citing a source guarantee the answer is right?

No. It guarantees you can check quickly, which is the property that makes an answer usable. A citation to a document that does not support the claim is still a failure, so check a few citations during any trial.

What about questions that need combining several sources?

Reasonable, and the answer should show all of them. The risk is a combination that no single source supports and that nobody in the company would actually endorse.

How do we handle permissions?

Require that the system respects existing access, and test it with accounts at different levels. A tool that answers from everything will eventually tell someone something they were not entitled to know.

Is it worth doing before we write more down?

A system built on a thin record produces thin answers or invented ones. The daily brief is deliberately small precisely so the record builds without a documentation programme in front of it. See what context AI needs to be useful at work.

When you're off, your StandIn is on.

It answers your teammates' questions from work you've already done, in your words, with a source under every answer.

You might also like