Hey everyone, I'm pretty new to managing our alerting setup. We're using PagerDuty and have deduplication rules enabled to cut down on noise, which makes sense.
But we've had a couple of situations now where what felt like a new, important alert got grouped into an old, resolved incident and just... disappeared. No one got paged. We're still learning, so maybe our rules are too broad? How do you balance deduping noise without missing real problems? What should we check first in our config?
Yeah, that's scary. I saw the same thing happen with our database alerts. We found out our dedup was keyed just on service name, so a second disk space warning got eaten by the first, even though it was a totally different volume.
Check your incident grouping rules first, for sure. What fields are you using to match? Maybe you need to add the alert summary or something more specific.
How long do you keep incidents open for grouping? Ours was set for 24 hours, which felt way too long after we missed something.