Every backup job over NFS fails all morning, and the root cause turns out to be the file service on the appliance, which stopped quietly the night before. Hardware alerting had been firing healthily all month. What does the monitoring configuration lack?
Select an answer to reveal the explanation.
Short Explanation
Hardware perfectly healthy and backups completely dead at the same time - that's the classic blind spot: nobody was watching the services themselves. Protocols like NFS and CIFS are daemons that can stop for their own reasons, and 'stopped' is a state your monitoring should know. Watch what your workload actually touches.
Full Explanation
DD OS file and protocol services can be enabled, disabled, or stopped independently of hardware state, so service liveness is a separate monitoring dimension from sensors and capacity. Polling that dimension means checking each required service is enabled and answering, so a stopped NFS daemon becomes an overnight alert with remediation runway instead of a morning of failed jobs - and it names the service instead of leaving consumers to infer from timeouts. Declaring service state out of scope abandons a first-class failure mode at the layer that owns it: consumers cannot distinguish a stopped service from an application bug until failures pile up. Attributing the stop to hardware is unsupported - daemons fail from configuration changes, upgrade side effects, and resource conditions with clean sensor readings - so faster hardware polling watches the wrong dial. A backup-application watchdog restarting storage services inverts the control relationship: a client should not command storage daemons, auto-restart masks the state change needing diagnosis, and the silent-stop pattern recurs anyway. Exam caveat: distinguish service-down from service-unreachable - pair the endpoint state check with a reachability probe from the client side so daemon faults and network faults separate cleanly. Operational check: during a window, stop a non-critical test service such as FTP on the appliance and confirm an alert naming that service reaches the on-call channel.