Runbook

Failure checks, queue replay and system handover.

Sample document based on a reference build. Real projects get the same structure.

What the system does

The lead-intake reference build accepts enquiries, normalizes them, qualifies the request, creates a CRM deal and notifies the owner. Uncertain requests stay available for manual review. The reply can be automatic or require approval, as agreed before launch.

Flow

Intake → Normalize → Qualify → Confidence check → CRM → Reply → Slack
                       ↓ uncertain              ↓ unavailable
                    Manual review            Queue → Replay

Access and retention

The model key and CRM token are stored in n8n credentials. They must not be pasted into node text, logs or this document. The CRM credential is limited to the required object and create permission. The owner controls the accounts and administrator access.

The reference design retains enquiry text for incident review for 30 days by default, then deletes it. Agree a different period before using this structure with real personal data.

Responding to alerts

Open the failed execution and identify the last completed step. Inspect the original input, provider response, branch decision and CRM write result. Check provider availability and credential permissions before changing the workflow.

Model failures get three attempts with exponential delay. After the attempts are exhausted, the unmodified enquiry should be marked for review. If it is missing from review, investigate the routing before retrying the whole workflow.

An alert after 15 minutes without a processed enquiry during business hours calls for an intake check. Confirm that the upstream channel is receiving requests and that it reaches the workflow. A quiet channel can be legitimate; record that finding instead of disabling the alert without review.

Replaying the queue

Confirm that CRM access works again. Inspect the queued item’s original source, timestamp and idempotency key. Keep that key unchanged. Replay the queued write through the configured recovery path, then check the CRM record and notification result. A replay must not create another deal for the same event.

The exact queue controls depend on the deployment. This sample does not prescribe an unverified command or production queue name. The delivered runbook records those controls for the actual system.

Changing the confidence threshold

Record the current threshold and the reason for the change. Update the confidence branch in a test copy of the workflow, then rerun the agreed acceptance examples, including short and mixed-language enquiries. Inspect which items move between automatic processing and manual review. Keep the previous configuration available for rollback and record the approved setting in the project documentation.

Handover and support

Keep the process map, acceptance examples and walkthrough in the owner’s account alongside this runbook. Support: dmitry@orbient.pro. Maintenance for one workflow is USD 150 per month; work outside the agreed scope is USD 95 per hour. Agree any incident response arrangements in the project terms.

Reference build and failure handling.