A runbook only counts if every failure mode in it actually happened

Handover isn't documentation plus a call. It's access narrowed, our own access removed, and a runbook proven against failures that actually happened — not ones we imagined.

We write two documents at the end of every engagement: an architecture doc and a runbook. Only one of them gets tested by someone hitting production at 2am, and it isn't the architecture doc.

Our house rule for a runbook: every failure mode in it has to be a real, dated incident with evidence — not a hypothetical someone thought was plausible. A runbook filled with imagined failures documents nothing anyone will actually see.

What a usable runbook actually contains

Each entry follows the same shape: symptom, the exact command that detects it, the root cause, the fix, and a "verify recovered" step. That last step isn't optional and isn't loose — recovery has to be checked the same way the failure was originally caught. The exact test that went red has to go green again; the exact question that failed has to pass again. A fix confirmed by re-asking a similar-sounding question isn't confirmed at all.

Handover is not a documentation deliverable

Our own "Handover" phase is explicit about the order: runbooks and knowledge transfer, then IAM narrowed to what production actually needs, then our own access removed. That sequence matters — narrowing access isn't a formality after the docs are done, it's the point where a client's team stops depending on us to keep the system running.

The test of whether it actually happened

On the two agents we most recently handed over, one runbook documents nine failure modes and the other eleven — every one a real incident, not a guess. But the number that matters more than the count is who reads them back: we walk the client's operations team through the runbook first, and only mark it accepted once they've used it themselves and returned with corrections. At that point both runbooks were still draft, under active negotiation with the client's own team — accepted for what they document today, not frozen as a finished artifact.

What we leave behind, deliberately

A single escalation order, printed in the runbook itself: the runbook first, because it should already contain the fix; a named advisory contact second; a named business owner for anything that changes what a number means; a named approver for anything going back to production. None of those roles is us, by design.

All insights

Talk to us

Tell us what system the answer lives in and who needs it. We'll reply with a view on whether it's a two-week assessment, a five-week pilot, or something else.

akash@insightnext.tech

InsightNext on LinkedIn