Oh wow, this thread is a goldmine, and it's exactly what I was coming here to learn. I'm setting up our alert routing now, and seeing everyone's real world pitfalls is incredibly valuable.
I'm leaning hard into the tag-based routing with a fallback like `team=catchall-infra`. The idea of making the tag a pre-provisioning requirement via OPA policy is terrifying from a change management perspective for us, but the alternative seems to be a mess of stale tickets or ignored queues, which is worse. My question is about the rule logic itself: if you're using a `shared=true` tag for platform resources, does your first rule filter for `shared=true` and send those to a platform team, and then your subsequent rules filter for specific teams like `team=frontend`? That ordering feels critical, and the automated sorting by specificity score that user517 mentioned sounds like a lifesaver.
Also, the point about webhook maintenance is something I wouldn't have considered for months. Attaching a metadata tag to the alert rule itself for audits is a clever, if slightly obsessive, fix. I'm already thinking about how to version that in our Terraform alongside the rules. Do you find you need to audit those channel-webhook mappings monthly, or is it more of a quarterly check?
Let's unpack that concrete example, because that's where the theory meets the pavement. You *can* absolutely set up a Slack channel for Team-A's security group changes and a PagerDuty endpoint for Team-B's IAM issues. The console will happily let you do it.
The problem is the moment someone in Team-A creates a resource and forgets the tag, that alert doesn't vanish. It either goes to the default global rule (your SecOps team) or, if you're smart and set one up, to a catchall team. So you've traded a blunt instrument for a Swiss cheese sieve.
Everyone's fixating on tag enforcement, but they're glossing over the operational tax. Who's maintaining the map of 50 team names to their correct Slack webhooks? What happens when Team-B renames their PagerDuty service? The routing breaks silently. You'll see "action succeeded" in the logs while the actual team never gets the page.
— skeptical but fair
It's mostly tags and integrations, like everyone's saying. But the Event Sync to SIEM question is interesting. We send everything to Splunk too, but it's for archival and correlation, not primary alerting. The latency is too high for actionable alerts in a live incident.
Your concrete example is totally doable. You'd make two alert rules: one filtered for security group changes plus the `team-a` tag, action set to their Slack webhook. Another filtered for IAM/S3 events plus the `team-b` tag, action set to their PagerDuty service. The console supports that.
The catch is you'll need a third, catch-all rule with no team tag filter that routes to a central triage channel for untagged resources. Otherwise, those alerts just hit the default global rule.
Tags are a fool's errand unless you enforce them at creation. And even then, the alert rule logic itself becomes a full time job to maintain.
You're asking about patterns? The pattern is everything breaks. Account groups get you halfway, then you're stuck with a hundred granular rules that nobody wants to own. The SIEM route just moves the complexity downstream.
> Can you set up different channels per alert rule, or is that a pipe dream?
It's possible, technically. But the webhooks rot. Channel names change. Teams get renamed. You'll spend more cycles auditing and fixing your routing matrix than actually responding to alerts.
The blunt instrument is blunt for a reason.
-- old school
Yes, you can absolutely set up different Slack channels per rule, that's the easy part! The real trick is making sure your tags are reliable before you build that whole system.
Your concrete example is 100% doable. One alert rule for security group changes filtered by `team-a` tag, action to their Slack. Another rule for IAM/S3 filtered by `team-b`, action to their PagerDuty. You'll use account groups for the broad scope and tags for the fine slicing.
But listen, the Event Sync to Splunk is for history and forensics, not for day-to-day alert routing. The latency will kill you when you need a quick response. And the webhook maintenance everyone's warning about is real - we have a quarterly check to make sure our channels and endpoints still exist. It's a chore.
Start with tag enforcement in your pipelines, or at least a strict `team=catchall` default. Otherwise you're just building a beautiful maze that half the alerts will never run through!
Happy customers, happy life.