Internal operations documentation checklist. Store credentials separately and apply access controls appropriate to the systems described.
Stuart Kerr Spindlow has confirmed personal use and testing of the software covered by Happy SysAdm. The assessments here distinguish documented behaviour from measured results; worked scenarios are labelled and are not personal test records.Begin with the service map
Prefer a controlled runbook linked to authoritative asset records over a collection of copied specifications. The runbook should explain the normal state, the permitted change and the expected outcome; the inventory should retain identifiers and ownership. This division avoids contradictory copies while giving a new responder enough context to act. Sensitive recovery instructions need a usable access path as well as restrictions.
Name the business service and its owner, then list infrastructure, SaaS dependencies, network paths and support arrangements. Record identifiers that survive a change in display name. Include where the authoritative inventory lives so the document does not become a conflicting second database.
Describe the normal state and what users depend on. A list of server specifications is useful but does not tell the next engineer which application fails when a host is unavailable.
Organise documentation around a service
- Service record
Business owner, technical owner and escalation route.
- Linked operating records
Architecture and asset identifiers, routine procedures, recovery runbooks and access-request process.
- Document control
Version, review owner, supported scope, review date and permitted audience.
Reference the approved secret store; never paste credentials into the map.
Make access and procedures usable
Record the roles required, the request and approval process and the emergency access route. Reference the secret store without copying credentials into pages or diagrams. Keep sensitive architecture and recovery information accessible only to people who need it.
For each routine task, give prerequisites, steps, expected results and a failure path. Include backup, restore, patching, renewal and offboarding. A command without its scope or expected result is difficult to use safely during an incident.
Use a controlled document record
Add an owner, last substantive review date, supported versions and the next review trigger. Link the runbook to the relevant asset and service records. Preserve revision history and make obsolete documents clearly non-authoritative.
The checklist below is suitable for a handover review. Ask the recipient to find and interpret each item. Testing discoverability is valuable: a complete document that nobody can find under pressure is still an operational gap.
| Area | Minimum useful record |
|---|---|
| Ownership | Business owner, technical owner and escalation |
| Architecture | Dependencies, identifiers and network boundaries |
| Operations | Routine procedures and expected results |
| Recovery | Recovery points, keys process and tested runbook |
| Governance | Version, review owner, date and access scope |
Prove the handover
Have an authorised colleague perform a low-risk task using only the documentation. Record ambiguous steps, missing permissions and undocumented dependencies. For recovery procedures, use a controlled exercise with isolation and business acceptance.
Keep an independent approved copy of critical recovery instructions if the normal documentation platform depends on the systems being restored. Review exports and backups of the documentation platform itself. Before changing tools, confirm that structure, attachments and access controls can be preserved or reconstructed.
Measure usable coverage
In a hypothetical set of 20 critical services, 16 current runbooks give 80% document coverage. If only 12 also have a verified independent recovery copy, coverage for that stricter outage requirement is 60%. Counting all pages in the wiki would not reveal this difference.
A handover check should record the precise task and the help required. Finding a page is different from completing its procedure unaided. No supplied handover record establishes a time saving here, so the recommendation is based on traceable ownership and recoverable access rather than a promised reduction in support time.