Conversational AI

The bot answers half the questions and honestly hands over the rest

Knowledge-base support that shows its source, says "I don't know" instead of inventing, and hands complex cases to a person.

Task
Knowledge-base first-line support
Timeline
5 working days
Stack
OpenAI · RAG · knowledge base · n8n
Result
2.6 s to a sourced answer

The problem

First-line support answers the same questions: where an order is, how to pay, whether an item can be returned and whether it is in stock. The queue builds up at night and operators clear it in the morning. Every interaction costs money, and roughly half are repetitive.

External benchmarks:

  • The median interaction involving a person is $13.50, based on Gartner customer-service benchmarks. By channel: phone $17–25, chat $10–16 and email $8–15. An AI interaction is around $0.50 versus roughly $6 for a person when the request type is suitable. These are external benchmarks, not measurements from this build.
  • Klarna, February 2024, published with OpenAI: in its first month the assistant handled 2.3 million conversations, two thirds of support chats, equivalent to the work of 700 employees. Resolution time fell from 11 minutes to under two, and repeat enquiries fell 25%. These are external company figures.
  • The required continuation: in May 2025 Klarna publicly reversed part of its AI strategy and began bringing people back. The CEO said the company had gone too far and focused too heavily on cost, hurting quality. The resulting model is hybrid: people use AI tools in each conversation.

The goal of first-line automation is not to replace people. It is to remove repetitive work and avoid lying about everything else.

What was built

  1. Knowledge base as the source of truth — documents are split into topic blocks for delivery, payment, returns, catalogue and order status, each with an identifier.
  2. Retrieval before generation — the model answers only from retrieved documents.
  3. Answer with a source — the interface highlights the document behind the answer. Both the user and operator can see the basis of the reply.
  4. Insufficiency flag — when the knowledge base has no answer, the model must return insufficient instead of inventing. The flow then hands the request to a person.
  5. Honest handoff — the interface states that a person and the agreed channel take over.
  6. Context limit — only the last 10 messages go to the model, keeping cost predictable and preventing the conversation from drifting.

Measured on the reference build on 2026-09-06: the answer to “Where is my order?” took 2.6 seconds, HTTP 200, and cited the document used. It correctly did not invent an order status and requested an order number in the NV-XXXXX format. Single measured run on 2026-09-06.

What happens when something breaks

  • The model is unavailable — the user sees a promise of a human response instead of an error, and the request enters the operator queue.
  • An answer is not supported by a document — it is not published. “I will check with a colleague” plus handoff is safer than a confident invention about return timing.
  • The knowledge base is stale — questions that trigger insufficient most often are collected into a report. This produces a direct list of missing content and makes improvement part of the system.
  • No actions involving money or orders — the bot does not cancel, refund or modify orders. It answers and hands off. Actions require a separate project with separate permissions.

Results

Metric Before After Nature of the figure
Time to first reply Minutes at peak, hours at night Seconds measured: 2.6 s
Requests resolved without a person 0% Planned 40–60% planned range; basis below
Cost per interaction $8–15 for email, $10–16 for chat ~$0.50 for automated requests Gartner benchmarks + model pricing
Repeat requests Baseline Reference −25% external Klarna data, not a guarantee
Answers without a source Not controlled 0 by construction architecture property

How this was calculated

Requests per month               = 1,200
Planned automated resolution     = 45% → 540 requests
Human handling time per request  = 6 min
Time returned                    = 540 × 6 ÷ 60 = 54 h/month
Australian support cost ≈ AUD 42/h fully loaded (≈ USD 27/h)
Staff time value                 ≈ 54 × USD 27 ≈ USD 1,458/month
Model + infrastructure           ≈ 540 × USD 0.05 ≈ USD 27/month
Monthly value                    ≈ USD 1,431/month (≈ AUD 2,200)
Build payback                    = USD 1,190 ÷ USD 1,431 ≈ 0.8 months
With maintenance                 = USD 1,190 ÷ USD 1,281 ≈ 0.9 months

Conservative case, 35% automation:
420 requests × 6 min = 42 h × USD 27 − USD 21 ≈ USD 1,113/month
Payback ≈ 1.1 months (≈ 1.2 months with maintenance)

Assumptions: 1,200 requests a month, 6 minutes of handling per request, a support specialist at AUD 42/h including super and on-costs, 1 AUD = 0.65 USD. The 45% resolution rate is a target, not a measured result.

Why 40–60%, rather than 76%. Vendors report 67–76% automated resolution. Independent production measurements report 38–53%, and B2B topics trail vendor figures by 17–25 percentage points. The case therefore uses an honest planned range of 40–60% for B2C with a good knowledge base. The eventual result depends on the request mix and knowledge quality, not on model choice. Promising 76% is a reliable way to fail acceptance.

What did not work on the first attempt

  1. The “entire knowledge base in the prompt” version answered confidently and incorrectly about availability and delivery timing. It was rebuilt to answer only from retrieved documents and to return an explicit “I don’t know” flag.
  2. Honest insufficient initially triggered too rarely because the model preferred to guess. The instruction had to reward refusal explicitly, and tests had to focus on questions outside the knowledge base.
  3. The first version exposed technical document IDs as sources. People did not understand them, so they were replaced with human-readable section names.

Timeline and cost

Day Work Deliverable
1 Collect three months of real questions, group topics, define acceptance Topic list and target automation rate
2 Structure, split and identify knowledge-base content Verifiable knowledge base
3 Retrieval, grounded generation and insufficiency flag Bot on real questions
4 Human handoff, interface and missing-content report Complete escalation flow
5 Run on real conversations, documentation and video walkthrough System and update rules in the owner’s account

External validation

Try it

Try the chatbot demo

Do you have a similar process?

Describe it in a message and within 48 hours I will tell you what is worth automating.

Describe the process