Week one of the new monitoring standard leaves the operations team drowning: the appliance was subscribed to 'everything', roughly sixty alert types arrive, and most have no owner or runbook. How should the subscription be fixed?
Select an answer to reveal the explanation.
Short Explanation
Sixty alert types with no owner isn't monitoring, it's a siren you've learned to sleep through. Curate the subscription to the failures this platform actually produces - drive and enclosure faults, capacity, temperature - and let the rest stay in the logs. 'Alert on everything' is what people say before they learn to read.
Full Explanation
Alert subscription is a design decision, not a volume dial. The working subscription mirrors the platform's genuine failure modes against the people who can act: drive and enclosure faults, capacity thresholds, temperature and power events - each subscribed because a named owner holds a runbook for it. Everything else stays in the logs - available for troubleshooting without shouting when nothing is wrong. Splitting the sixty across two teams relocates the drowning without subtracting one low-value alert - volume is the disease, not the org chart - and a second drowning team is one more owner lost. Muting everything for a calibration month inverts the risk: a newly handed-over system is exactly when hardware infant mortality and configuration mistakes surface; a quiet month buys tidy thresholds at the price of blindness in the highest-risk window. Raising everything to critical is noise laundering: priority exists so humans respond in proportion; where everything pages, nothing does, and deduplication collapses repetitions of one storm, not sixty distinct low-value event types. Exam angle: 'monitor everything' is an intake requirement; 'alert on what an owner can act on' is the deliverable. Operational check: every live alert category has a named owner and a runbook entry; categories missing both go back to log-only.