A municipal inspections platform is mostly asynchronous, yet reliability dashboards still alert only on HTTP 5xx from the web tier. Backlogs and poison messages grow unnoticed. Which observability improvement best fits async reliability?
Select an answer to reveal the explanation.
Short Explanation
Async systems fail quietly—the website can smile while the work pile rots. Watch queue depth, message age, and DLQs like you watch 5xx, or inspections will stall without a red banner.
Full Explanation
Reliability observability for asynchronous civic platforms must include queue backlog, message age, and dead-letter rates, not only HTTP 5xx. Those signals detect stalled workers and poison messages early. Removing DLQs or ignoring queue metrics blinds operators; idle CPU can indicate starved consumers rather than health.