The DR runbook promises recovery from a replica that is never more than an hour behind, yet no metric anywhere tracks how current the replication actually is. What monitoring configuration closes that gap?
Select an answer to reveal the explanation.
Short Explanation
An RPO you don't measure is a wish. Replication currency - how far behind the copy actually runs - is a metric like any other: set a threshold, raise an alert, put it on the dashboard. The day you learn the lag is nine hours should be a drill, not a disaster.
Full Explanation
Recovery promises live or die on data currency, so the currency figure the platform already tracks - last update per replication pair - should feed the monitoring system with a threshold that mirrors the stated RPO. When lag approaches the promise, operations gets time to fix transport or workload pressure, and the alert sits beside hardware and capacity alerts so one workflow owns appliance health end to end. Byte-count equality between sites proves nothing about timing: a replica can match totals today and still have fallen six hours behind last weekend while the comparison sat idle, and legitimately differing data makes equality noisy and untrustworthy. Schedule-enabled is a configuration echo, not an outcome check - pairs pause after errors, queues back up, and links degrade with the schedule looking perfectly healthy the entire time. Link utilization is a correlate, not the promise: a moderately busy link can stall behind a retry storm while a saturated one keeps currency comfortably, and the runbook defines its hour in staleness, not in percent-used. Exam caveat: currency is tracked per pair and per direction - configure the threshold on the side your recovery plan actually reads. Operational check: in a maintenance window, pause one test pair and confirm the alert fires within the threshold window and reaches the same people and tools as hardware alerts.