Adaptable recovery template for a small IT service. Fill it with verified internal details; the illustrative roles are not contact information.
Stuart Kerr Spindlow has confirmed personal use and testing of the software covered by Happy SysAdm. The assessments here distinguish documented behaviour from measured results; worked scenarios are labelled and are not personal test records.Make the first page useful during an outage
A dependency-ordered runbook is more useful than a long checklist organised by product. It explains why a responder must recover identity or storage before starting an application and which tasks another responder can do in parallel. Pair it with a short incident-facing summary and a controlled independent copy. A document that is available only through the failed service cannot be the sole recovery instruction.
Record the service name, runbook owner, version, review date and activation authority. Define the conditions that trigger recovery and who can declare an incident. Include a contact tree with primary and alternate roles, supplier contract references and an out-of-band communication channel.
Keep secrets out of the runbook. Reference the approved emergency credential process and verify that authorised responders can use it when normal identity services are unavailable. An inaccessible document repository can turn a good runbook into a missing dependency.
Recover in dependency order
Map the service to identity, network, DNS, storage, databases and external integrations. Restore the prerequisites before the application that consumes them. Name the exact supported procedure for each component and its expected output.
Use the sequence table as a template. Add the real system identifiers, acceptance checks and stop conditions. If a step fails, the next action should say whether to retry, roll back, select another recovery point or escalate. Avoid open-ended instructions such as fix networking.
| Stage | Owner role | Gate before continuing |
|---|---|---|
| Declare and contain | Incident lead | Scope, authority and communication agreed |
| Recover prerequisites | Infrastructure lead | Identity, DNS, storage and access verified |
| Restore data | Recovery operator | Correct point and destination confirmed |
| Validate service | Application owner | Business transaction accepted |
| Resume operation | Incident lead | Monitoring, communications and rollback ready |
Recover the prerequisites before the service
- Prerequisites
Identity, network, DNS, storage and emergency access; use the service-specific dependency order.
- Data and application
Choose an authorised recovery point and restore through supported procedures.
- Business acceptance
Validate a real transaction, permissions and monitoring before traffic resumes.
Stop at a failed dependency. Production overwrite, traffic switching and failback require explicit decision gates.
Put decision gates around dangerous transitions
Before restoring over production or switching traffic, confirm the recovery point, current incident containment and the effect on data written since that point. Get the required business decision on data loss. Preserve evidence where an incident investigation is under way.
Define a return-to-primary plan separately from initial recovery. Moving traffic back may need reconciliation or another outage. Record who can approve failback and which measurements show that it is safe.
Exercise the document, not just the infrastructure
Ask a responder who did not write the runbook to follow it in a controlled exercise. Record every missing permission, ambiguous instruction and dependency that required outside knowledge. Measure the full recovery path against the service objective.
After the exercise, update the sequence and verify access to the revised copy. Set a review trigger for infrastructure changes and personnel changes, as well as a periodic date. Keep an offline or otherwise independent controlled copy and retire stale copies deliberately.
A completed timing example
Consider a hypothetical sequence of 15 minutes to declare and contain, 30 for prerequisites, 90 for data recovery and 25 for service acceptance. Its sequential critical path is 160 minutes. A three-hour target leaves 20 minutes for delays; a 30-minute supplier wait would push the same plan to 190 minutes and miss the target by ten minutes.
The conclusion is to remove or pre-arrange that supplier dependency where feasible, rather than silently shorten the acceptance step. These are planning calculations, not an exercise result. The measured run should retain both the task timeline and the final acceptance time so a revised plan addresses the actual delay.