Why a support bot must be able to say “I don’t know”

Retrieval, an explicit insufficient flag and a handoff that keeps the request moving.

The first version of my support reference build put the whole knowledge base into the prompt. It could produce a fluent answer about availability or delivery even when the supplied material did not establish the answer. The response sounded usable. That was the problem: someone reading it had no reliable way to tell where the documentation ended and the model’s guess began.

This is a write-up of that reference build, not a client deployment. The useful change was to make missing information an explicit outcome of the workflow. The system retrieves relevant documents, asks the model to answer from those documents, and routes the request to a person when they are insufficient. The interface has to explain that route as clearly as it explains a successful answer.

A policy cannot tell you where an order is

“Where is my order?” looks like a familiar support question. A delivery policy can explain the normal process, the information needed to look up an order and what a customer should do next. It cannot establish that a particular parcel has left the warehouse. That requires a source containing the particular order’s current state.

In the measured reference run on 6 September 2026, the bot requested an order number in the NV-XXXXX format and cited the document it used. It did not invent an order status. The response took 2.6 seconds in that single run. That timing tells us about one execution; it does not establish a general latency guarantee or prove that every possible question is handled correctly.

For this question, asking for the missing identifier was a useful answer. If the next step still cannot be supported by the available documents or permitted integrations, the request needs a person. Continuing the conversation with increasingly confident wording would not provide the missing order data.

Make uncertainty a branch

An instruction saying “do not make things up” leaves an operational question unanswered: what should happen when there is no supported answer? In this build, the model returns an explicit insufficient flag. The workflow can use that value to choose a handoff instead of publishing a speculative reply.

That flag initially appeared too rarely. The model preferred to complete the answer. I changed the instruction to explicitly reward refusal when the source material was inadequate, then tested questions outside the knowledge base. Testing only questions with obvious answers would have missed the behaviour that needed fixing.

The source requirement applies to the answer itself. A retrieved document can share words with a question without supporting the proposed response. A returns policy, for example, gives a policy; it does not prove that a specific return has been accepted. During review, I look at the claim and the retrieved passage together. A nearby topic is not enough evidence for a specific commitment.

The handoff needs a destination

A refusal without a next step leaves the customer to start again. The reference flow tells the user that a person will take over through the agreed channel and puts the request into the operator queue. The same path is needed when the model provider is unavailable. Customers should not need to interpret a provider error to understand what happens to their question.

This bot answers and hands off. It does not cancel orders, issue refunds or change order records. Adding those actions would require a separate scope and permissions. Keeping that boundary visible makes it possible to review the answering behaviour without assuming that a plausible sentence authorizes a transaction.

Sources should mean something to the reader

The early interface displayed technical document identifiers. Those identifiers were useful internally, but customers could not tell what they referred to. I replaced them with readable section names. A source label should help someone find the basis for an answer; exposing an internal ID alone did not achieve that.

Repeated insufficient questions also provide a practical maintenance list. They show where the knowledge base is missing content or where a process needs a clearer handoff. The report does not automatically turn those questions into approved answers. Someone still has to supply and check the underlying information before the bot can use it.

The support reference case contains the flow, failure handling and the distinction between measured response time and planned resolution rate. Its acceptance work includes questions the documents cannot answer, because those questions exercise the branch that a normal happy-path demo leaves unseen.