Scope & evidence

Storage planning concepts for small IT environments. Failure tolerance and rebuild behaviour depend on the exact RAID level, controller and drive configuration.

Stuart Kerr Spindlow has confirmed personal use and testing of the software covered by Happy SysAdm. The assessments here distinguish documented behaviour from measured results; worked scenarios are labelled and are not personal test records.

Name the failure you are protecting against

Use RAID to address the availability consequences of supported drive failures and backups to recover an earlier usable state after deletion, corruption or wider loss. A mirrored current copy faithfully reproduces many unwanted changes, so it cannot replace retained history. For a service that needs both continued operation and recoverable data, the appropriate decision is usually a layered design rather than selecting one technology as the winner.

A redundant array may continue operating when a member drive fails within its tolerance. The same array can faithfully apply an accidental deletion or encrypted overwrite across its members. A second copy of the current state is not necessarily a recoverable older state.

List drive failure, controller failure, operator error, malware, theft and site loss separately. Map each to a protection mechanism and an actual recovery procedure. Avoid describing one technology as protection from every failure category.

Decision map

Availability and recovery solve different failures

Member-drive failure
A redundant RAID level may maintain service within its tolerance; degraded operation still needs attention.
Deletion or corruption
The array may apply the unwanted change. Recovery requires a usable retained point.
Whole-system loss
Recovery needs a copy, credentials and keys that survive the source failure domain.

RAID 0 is not redundant. No RAID layout is an independent retained backup by itself.

Failure-scenario comparison. RAID level and backup separation determine the protection actually available. Evidence sources.

Account for degraded operation

When an array is degraded, remaining drives and rebuild work can affect performance and exposure. Monitor controller health and replacement status. Confirm the supported replacement procedure and identify the correct physical drive before removing anything.

Do not infer fault tolerance from raw drive count. RAID levels differ, and multiple failures can have different consequences depending on placement and state. Keep the controller configuration and vendor support information available without relying on the failed array.

Build an independent recovery path

Keep backups with appropriate retention and separation from the source system’s failure and access domains. Protect repository administration and decryption material. Test recovery of files and the application, including permissions and dependencies.

A local snapshot can complement this design but may still share the same chassis, controller or credentials. Decide whether an off-site or immutable copy is needed for the incident scenarios you identified. The useful question is what survives, not how many storage features are enabled.

Verify both availability and recovery

Test monitoring and the replacement workflow through a supported exercise or maintenance process. Do not deliberately pull production drives merely to demonstrate redundancy. For backups, restore representative data into isolation and obtain acceptance.

Document the residual risks and expected outage. RAID can reduce one kind of interruption while backups recover data after another. Budget and operate both according to the service’s requirements rather than treating them as interchangeable purchases.

Capacity does not measure backup protection

As a simple capacity illustration, two equal 4 TB drives in a basic two-way mirror provide about 4 TB before formatting and reserves, not 8 TB of independent historical recovery. There are two current copies of blocks, but a propagated deletion can affect both.

The capacity arithmetic is not a fault-tolerance guarantee for every controller or array state. Restoration from a retained independent copy is a separate acceptance case from continuing to run after a member failure. A successful rebuild does not prove that an older version of a deleted file is recoverable.

References

Next useful steps

Read our editorial and corrections policy.