Scope & evidence

Adaptable recovery template for a small IT service. Fill it with verified internal details; the illustrative roles are not contact information.

Stuart Kerr Spindlow has confirmed personal use and testing of the software covered by Happy SysAdm. The assessments here distinguish documented behaviour from measured results; worked scenarios are labelled and are not personal test records.

Make the first page useful during an outage

A dependency-ordered runbook is more useful than a long checklist organised by product. It explains why a responder must recover identity or storage before starting an application and which tasks another responder can do in parallel. Pair it with a short incident-facing summary and a controlled independent copy. A document that is available only through the failed service cannot be the sole recovery instruction.

Record the service name, runbook owner, version, review date and activation authority. Define the conditions that trigger recovery and who can declare an incident. Include a contact tree with primary and alternate roles, supplier contract references and an out-of-band communication channel.

Keep secrets out of the runbook. Reference the approved emergency credential process and verify that authorised responders can use it when normal identity services are unavailable. An inaccessible document repository can turn a good runbook into a missing dependency.

Recover in dependency order

Map the service to identity, network, DNS, storage, databases and external integrations. Restore the prerequisites before the application that consumes them. Name the exact supported procedure for each component and its expected output.

Use the sequence table as a template. Add the real system identifiers, acceptance checks and stop conditions. If a step fails, the next action should say whether to retry, roll back, select another recovery point or escalate. Avoid open-ended instructions such as fix networking.

StageOwner roleGate before continuing
Declare and containIncident leadScope, authority and communication agreed
Recover prerequisitesInfrastructure leadIdentity, DNS, storage and access verified
Restore dataRecovery operatorCorrect point and destination confirmed
Validate serviceApplication ownerBusiness transaction accepted
Resume operationIncident leadMonitoring, communications and rollback ready
Dependency map

Recover the prerequisites before the service

  1. Prerequisites

    Identity, network, DNS, storage and emergency access; use the service-specific dependency order.

  2. Data and application

    Choose an authorised recovery point and restore through supported procedures.

  3. Business acceptance

    Validate a real transaction, permissions and monitoring before traffic resumes.

Stop at a failed dependency. Production overwrite, traffic switching and failback require explicit decision gates.

Illustrative dependency layers. This is not a universal order within each layer. Evidence sources.

Put decision gates around dangerous transitions

Before restoring over production or switching traffic, confirm the recovery point, current incident containment and the effect on data written since that point. Get the required business decision on data loss. Preserve evidence where an incident investigation is under way.

Define a return-to-primary plan separately from initial recovery. Moving traffic back may need reconciliation or another outage. Record who can approve failback and which measurements show that it is safe.

Exercise the document, not just the infrastructure

Ask a responder who did not write the runbook to follow it in a controlled exercise. Record every missing permission, ambiguous instruction and dependency that required outside knowledge. Measure the full recovery path against the service objective.

After the exercise, update the sequence and verify access to the revised copy. Set a review trigger for infrastructure changes and personnel changes, as well as a periodic date. Keep an offline or otherwise independent controlled copy and retire stale copies deliberately.

A completed timing example

Consider a hypothetical sequence of 15 minutes to declare and contain, 30 for prerequisites, 90 for data recovery and 25 for service acceptance. Its sequential critical path is 160 minutes. A three-hour target leaves 20 minutes for delays; a 30-minute supplier wait would push the same plan to 190 minutes and miss the target by ten minutes.

The conclusion is to remove or pre-arrange that supplier dependency where feasible, rather than silently shorten the acceptance step. These are planning calculations, not an exercise result. The measured run should retain both the task timeline and the final acceptance time so a revised plan addresses the actual delay.

References

Next useful steps

Read our editorial and corrections policy.