Did the Alert Recover by Itself, or Was It Closed by a Rule? Reviews Cannot Treat Them as One State
A Review Gets Stuck on One Question
Twenty minutes before the morning meeting, operations lead Xiao Zhou is asked a very specific question: for yesterday afternoon's batch of API timeout alerts, did the service recover, or were the alerts automatically closed by a rule?
He has the alert list and the closure rate. The red dots have disappeared from the list, and several alerts are no longer active. But as the team keeps tracing downward, the room gets stuck: some alerts were closed by the on-call engineer, some were pushed back by recovery events, and a few aggregated alerts were automatically closed after passing the inactivity window.
At this point, "they are all ended" becomes the least useful answer.
What an alert review really needs to trace is not whether the red dot is still there, but why it left the scene.
