Scope & evidence

Vendor-neutral checklist for supported server operating systems and applications. Thresholds must be calibrated to the workload; no universal CPU or disk threshold is implied.

Stuart Kerr Spindlow has confirmed personal use and testing of the software covered by Happy SysAdm. The assessments here distinguish documented behaviour from measured results; worked scenarios are labelled and are not personal test records.

Start at the service boundary

A service check paired with a small set of explanatory host metrics is a better starting point than collecting every available counter. Google’s SRE guidance groups user-facing monitoring around latency, traffic, errors and saturation. Those signals help distinguish an unavailable service from an idle machine; host CPU and memory explain symptoms but should not stand in for the user operation.

List the operations users depend on and how you can check them safely. A health endpoint, a synthetic transaction or a supported database probe gives different evidence from a ping. Decide what successful behaviour looks like and which failures require immediate action.

Record service ownership, maintenance windows and dependencies. A network outage can make many hosts appear down; avoid treating every symptom as an independent incident. Monitor the monitoring system and its notification delivery too.

Dependency map

Start with the operation the user needs

  1. User operation

    Can the transaction complete? Observe errors and response time.

  2. Service dependencies

    Check the application and the services it depends on.

  3. Host and collection

    Use CPU, memory, storage and network evidence; detect missing telemetry too.

Each actionable alert also needs an owner, response and tested delivery path.

Investigation layers. A responsive host does not establish a working application. Evidence sources.

Collect a small useful baseline

Capture availability, request volume, errors and latency where the application exposes them. Add host CPU, memory pressure, storage latency, free capacity and network errors to explain symptoms. Preserve enough history to compare normal working periods, scheduled jobs and unusual load.

A single average can hide a short but serious delay. Choose an aggregation and time window that match the user impact. Do not collect detailed logs indefinitely without considering sensitivity, storage cost and access control.

Attach a response to each alert

For each proposed notification, write what the responder should do and how quickly. An alert that nobody can act on may belong in a report instead. Use persistence windows or recovery thresholds where brief fluctuations would otherwise create noise.

The worksheet below separates checks from responses. Fill in actual thresholds after observing the service. Include growth rate and lead time for capacity alerts: a volume with substantial free space may still need action if it is filling faster than capacity can be added.

SignalQuestionResponse record
AvailabilityCan the user operation complete?Owner, dependency checks and escalation
Latency/errorsIs service quality degrading?Time window and diagnostic evidence
CapacityWhen will headroom be exhausted?Growth rate and expansion lead time
CollectionAre metrics or checks missing?Agent, credentials and monitoring health

Test the complete route

Use a controlled test to verify collection, rule evaluation, delivery, acknowledgement and escalation. Do not create an outage merely to prove a notification. A supported test event or isolated service can validate much of the route.

Review alerts after incidents and changes. Remove obsolete hosts, correct owners and check credentials before they expire. Keep dashboards for investigation, but make sure urgent conditions reach a person without requiring somebody to watch a screen continuously.

Why the average can hide a bad experience

In a hypothetical sample of 100 requests, 95 take 100 ms and five take 2,000 ms. The mean is (95 × 100 + 5 × 2,000) ÷ 100 = 195 ms, although five users waited two seconds. An average below 200 ms therefore would not establish that every request met a 200 ms objective.

With the nearest-rank percentile convention, p95 in this example is 100 ms and p99 is 2,000 ms. State the aggregation method and sample window: a percentile without them can be misleading. These calculated examples explain metric choice; they are not latency results from a monitored service.

Give storage alerts a measurable response window

For each important volume, record free bytes and percentage alongside an operational reserve. Estimate time to reserve using representative positive growth, then compare that with the time needed to investigate and complete an approved change. Keep absolute-space and missing-data conditions separate so a stale metric or a rapidly growing volume is not hidden by a percentage rule.

A useful acceptance check covers the intended volume, the rule, notification delivery, ownership and recovery. Use a safe test condition or isolated lab rather than filling production storage. Record warning and escalation conditions explicitly, including whether combined rules use AND or OR.

References

Next useful steps

Read our editorial and corrections policy.