A flapping network interface generates two hundred overnight pages and buries one genuine disk-failure alert in the flood, which nobody reads until morning. What configuration thinking would have kept the meaningful event visible through the storm?
Select an answer to reveal the explanation.
Short Explanation
An alarm that cries wolf trains you to sleep through it. Configure severities so only actionable events page you, and add rate control so one noisy source can't flood the line. The disk alert was actually there - the configuration is what made it invisible.
Full Explanation
Usable alerting treats the notification channel as a scarce resource and engineers it that way. Alerts should be matched to response - what pages, what tickets, what only logs - and rate controls or aggregation collapse repeats from one source into a single actionable event. Applied here, the flapping link produces one page with context, the disk failure still pages above it, and the on-call human sees signal instead of noise. Time-window muting is dangerous rather than clever: failures do not consult calendars, and a disk dying at 03:00 during a mute window is exactly the loss the monitoring existed to prevent. Raising everything to critical destroys the severity mapping itself - when all events are urgent, the operator cannot triage at all, and the two-hundred-page night returns unchanged under a redder label. Splitting event classes across teams reduces neither volume nor blindness; the network queue drowns in flaps just as happily, and correlated failures routinely span both domains, so a split inbox adds handoff latency to the burial. Exam caveat: rate control must collapse repeats without suppressing the first occurrence of a genuinely new condition, or throttling becomes its own blind spot. Operational check: in a maintenance window, induce a repeatable benign event and confirm a simultaneously raised second alert still arrives promptly and separately at the on-call channel.