Printable exercise plan for an authorised restore of a representative workload. Use product-specific restore documentation and an approved destination.
Stuart Kerr Spindlow has confirmed personal use and testing of the software covered by Happy SysAdm. The assessments here distinguish documented behaviour from measured results; worked scenarios are labelled and are not personal test records.Define the exercise and containment
An alternate-location restore is the preferred routine exercise because it can demonstrate usable recovery while preserving the source. A repository integrity check is valuable, but it cannot establish that the application starts, the right users retain access or the business transaction succeeds. The stronger acceptance test follows the recovered data all the way to the task the service exists to perform.
Choose a service, incident scenario and recovery point. Name the exercise lead and the person who can approve the recovered service. Set the expected RPO and RTO and describe what is excluded from the exercise.
Prepare a destination that cannot accidentally contact production. A recovered system may send email, run scheduled jobs, duplicate an address or connect to a live database. Check routing, DNS, integrations and credentials before starting it. Do not rely on a label saying Lab as proof of isolation.
Check prerequisites before the timer starts
Confirm access to the repository, decryption material, installation media and supported restore procedure. Document whether these were already available or had to be recovered. Check capacity for restored data, temporary files and logs.
Preserve the original source and backups. Do not overwrite production as a convenience for a routine exercise. If the product has an original-location restore option, review the selected destination and recovery mode with a second authorised person before committing.
Restore and validate the service
Record start times for retrieval, data restoration and application readiness. Verify data from the selected point, application consistency, permissions and a meaningful user transaction. For a file service, test representative files and access controls; for an application, test the business operation rather than only its login page.
Use the evidence log below. Keep screenshots or logs in a controlled evidence store, with secrets and personal information handled under the organisation’s rules. A checksum can help verify file integrity but does not prove that a multi-part application is operational.
| Check | Evidence to capture |
|---|---|
| Recovery point | Timestamp, consistency type and source job |
| Isolation | Network and integration checks |
| Application | Transaction performed and expected result |
| Permissions | Authorised and unauthorised access cases |
| Acceptance | Owner, outcome, elapsed time and remaining gaps |
Restoration finishes at service acceptance
- Contain
Verify isolation, destination and recovery-point selection.
- Restore
Record retrieval and restore times without overwriting production.
- Validate
Check data, permissions and a meaningful business transaction.
- Accept
Owner records pass, partial pass or fail; assign gaps and clean up.
Record whether prerequisites were already available when the exercise began.
Close the exercise and fix the gaps
Obtain an explicit pass, partial pass or fail from the service owner. Compare the measured result with the agreed objectives and record the limiting step. Assign remediation with an owner and due date.
Remove temporary access, stop restored scheduled tasks and securely dispose of test copies when the retention rule permits. Confirm that production and backup jobs remain healthy. Repeat the failed portion after remediation and keep both results so the improvement is traceable.
Separate file integrity from service acceptance
An AI-run check on 23 September 2026 used restic 0.19.1 on macOS 26.6.2 to restore an earlier snapshot after overwriting one synthetic file and deleting another. All 12 original files, totalling 12 MiB, returned with matching SHA-256 hashes, and the repository data-read check passed. This establishes byte equality for those fixtures, not a 100% future recovery success rate. The source and repository were on the same host; permissions, application transactions and off-site recovery were not tested.
Measure the recovery clock from the agreed incident or exercise start to accepted service, while recording prepared prerequisites separately. Pre-staging credentials and storage can be entirely sensible, but the result should not be reported as a cold recovery if those preparations were outside the measured interval.