Team size: 12 engineers, 3 SREs.
Stack: AWS (MSK, EKS), ClickHouse for analytics, Snowflake for warehousing. Security team uses Panther for SIEM, but alert noise is high.
We needed Claw (our internal threat detection tool) alerts to be actionable. Slack is our incident response hub. The bot filters severity >= HIGH and posts with context.
Key requirements:
* Low latency (<5s from alert to Slack).
* Structured message with one-click link to full alert in Claw UI.
* Avoid spamming the channel.
Solution: A small Go service subscribes to Claw's Kafka alert topic, applies filtering, formats, and calls Slack's Incoming Webhook.
Core filtering logic:
```go
if alert.Severity < config.ThresholdSeverity {
return // Drop
}
if slices.Contains(config.IgnoredRuleIDs, alert.RuleID) {
return // Drop
}
```
Message builder creates a Slack block with:
* Alert title & rule ID.
* Timestamp (converted to local TZ).
* Affected resource (e.g., `arn:aws:ec2:...`).
* Direct link button.
Considered self-hosting the bot (simple K8s deployment), but opted for a managed Fargate task due to team bandwidth. Runs for ~$12/month.
Results: Alert-to-acknowledgement time dropped from ~15 minutes (email) to under 2 minutes. Channel gets ~3-5 high-signal posts per day.
Numbers don't lie.
Curious about the Fargate choice here. At ~$12/month for a small Go service, I'm wondering if you considered Lambda with Kafka as an event source? The pricing would likely be under $1/month given the low volume, and you could still keep the latency under 5s. Was it about the Kafka library support or just keeping ops simple?
Lambda's Kafka event source wasn't ready when we built this. The Go library support is fine now, but you're right on cost.
We keep the alert consumer co-located with other services in the same ECS cluster. It simplifies our monitoring and networking. One less moving part.
Ship it, but test it first
> keeping ops simple
That's a fair point, but I'd be worried about cold starts with Lambda if the alert volume is sporadic. Even a two-second delay could feel critical for a HIGH severity security alert. Is that a real trade-off people see, or is it mostly theoretical now?