Alerts now tell you once per problem, not once per check
If you have ever muted a MetricsWatch alert because it kept emailing you about a problem you already knew about, this release is for you. An alert now tracks the problem, not the check: one email or Slack message when it starts, quiet while it continues, one when it clears.
What used to happen
MetricsWatch checked your metric on a schedule. If the condition was true, it sent you an email or Slack alert. If it was still true on the next check, it sent another, once per cooldown window, for as long as the problem lasted. A metric wobbling either side of your threshold was worse, because every wobble back over the line counted as a brand new problem.
What happens now
Your alert tracks an incident.
It tells you when the problem starts. The first check that crosses your threshold opens an incident and sends one notification. Every later check that is still over the line sends nothing.
It re-arms once things have settled. For an alert that checks in real time, the incident closes after three clear checks in a row. A wobble back over the line inside those three is still the same incident, not a new one.
Daily alerts work by the day. An alert that checks once a day tells you on every day the condition is met, and closes on the first clear day. At one check a day there is no noise to group, so we do not make you wait three clear days to hear that a problem has come back.
It ignores percentage swings on tiny numbers. If your alert is a percentage rule, we now need at least 20 of whatever you are measuring on the larger side before acting. Going from zero conversions to one is an infinite percentage increase, and we used to email you about it. A low-volume reading also does not count as recovery, so a metric dropping under that floor while an incident is open will not quietly close it.
"Resolved" only arrives if you heard about the problem. If your alert notifies during business hours only, and the condition tripped at 2am and cleared at 9am, the old rules sent a "Resolved" message about a problem you were never told about. Now a recovery message goes out only if the notification that opened the incident actually reached you.
The practical ceiling for a real-time alert: at most two messages per incident. One when it starts, one when it is over.
Does this mean fewer emails?
For most alerts, substantially. We replayed 30 days of real checks from the customer accounts that run alerts today through both sets of rules: 122 notifications under the old rules, 65 under the new ones. The alerts that fire most often gain the most.
For a small number of alerts it means slightly more, because you now get a "Resolved" message you never used to get. If you would rather not have those, tell us which alert and we will look at it with you; the answer may be a setting on that alert, or leaving your account on the old behaviour.
Slack
The same rule applies to both channels. An incident that opens sends one email or one Slack message, whichever you have set up on that alert, and the quiet period is the same for both.
How and when this reaches your account
We are switching it on for existing accounts one at a time, not all at once, and nothing changes on your alerts until it reaches your account.
If you would rather be notified on every check that meets an alert's condition, an account admin can open Settings, choose Notifications, find Grouped Alert Notifications, untick Group repeated alert notifications into one incident, and select Save. This control appears once grouped notifications have reached your account.
How to see it working
The easiest confirmation is the next time something goes wrong: one message when it starts, silence while it continues, one when it clears. In the alert's log, a check that was seen but deliberately not sent carries the reason it was held, so a quiet check is never a missed one.