Scope & evidence

Vendor-neutral diagnostic workflow for modern hypervisors. Use the platform’s supported counters and logs; this is not a replacement for legacy ESX-specific procedures.

Stuart Kerr Spindlow has confirmed personal use and testing of the software covered by Happy SysAdm. The assessments here distinguish documented behaviour from measured results; worked scenarios are labelled and are not personal test records.

Define the symptom and blast radius

Compare an affected VM with a suitable peer before adding CPU or memory. A single guest’s symptoms suggest a different starting point from several guests slowing on one datastore. Resource increases are justified when the relevant contention is measured; otherwise they can move the bottleneck or increase host pressure. The best first intervention is the one that tests a specific explanation while preserving the evidence.

Record the user-visible failure, start time and recent changes. Determine whether one VM, several VMs on a host or an entire service is affected. Check whether the console works when network access fails. A working console narrows the problem but does not establish that the guest application is healthy.

Compare with an unaffected VM using the same host, storage or network. Note which dependency is shared. Avoid rebooting immediately unless the incident plan requires it, because transient counters and logs may be lost.

Inspect the layers in a consistent order

Within the guest, check application logs, resource pressure and recent updates. At the virtual network layer, inspect adapter state, VLAN or port-group placement, addressing and DNS. At storage, compare latency, free capacity and errors across the guest and backing datastore.

On the host, inspect contention, hardware health and management events. A guest CPU percentage may not describe time spent waiting for host scheduling. Use the hypervisor’s own documented measurements rather than interpreting every counter as though it were a physical server metric.

Decision map

Use the blast radius to narrow the layer

One guest
Check the application, guest events and console versus network behaviour.
Shared network
Compare adapter placement, addressing, DNS and affected peers.
Shared storage
Correlate latency, free capacity and errors on the backing datastore.
Shared host
Inspect contention, hardware health and management events.

Preserve measurements, test one hypothesis at a time and re-run the original user operation.

Investigation map, not a diagnosis. More than one layer can contribute to a symptom. Evidence sources.

Change one variable with a recovery path

Form a testable hypothesis, such as a storage slowdown shared by VMs on one datastore. Capture a baseline and choose a low-risk check before adding resources or moving workloads. A migration can move the symptom while obscuring the cause.

Before resizing, changing controllers or restarting, confirm application impact, backup readiness and rollback limits. Do not delete snapshot or virtual-disk files manually to recover space. A storage emergency needs the platform’s supported consolidation or recovery procedure.

Verify the original workload

After a change, repeat the user operation and compare the same measurements over a representative period. Check other VMs sharing the affected resources. A fast login immediately after a reboot does not prove that a recurring workload problem is resolved.

Retain the timeline, relevant counters, platform versions and change outcome for escalation. If evidence points to hardware, storage or a vendor defect, stop speculative tuning and use the supported support path with the collected record.

Count the blast radius, then compare like workloads

If an illustrative host runs ten VMs and three affected VMs share one datastore, that is a useful grouping for investigation, not a 30% host-failure rate. The seven unaffected VMs are meaningful controls only if their workload, storage and observation periods are comparable.

For an illustrative operation that falls from 12 seconds to nine seconds under equivalent conditions, elapsed time falls by 25%. Its reciprocal throughput rises by about 33.3% only if repeated operations remain comparable. These calculated relationships are not measurements from this guide and do not establish a benefit from any particular VM setting.

References

Next useful steps

Read our editorial and corrections policy.