We log everything to a single CloudWatch group with a 1-day retention policy. That's enough to debug a recent alert without blowing up costs. For anything older, we rely on the structured data we forward to PagerDuty.
The cost wasn't bad for us, but we filter out debug-level logs in production to keep volume down.
Glad to see someone else taking the webhook route. That direct link to the investigation panel is the most important piece - it's the first thing our engineers need.
One thing to consider is structuring your Lambda to parse out the dataset slice or dimension from the payload. Arize often buries it in the metadata, but it's crucial context for whoever gets paged. We include it as a custom detail in PagerDuty so the on-call knows immediately if it's a data quality issue or specific to a region.
Your mention of Slack mirroring from PagerDuty is the right call. We found it's better to let PagerDuty handle the routing logic and just use Slack for visibility.
That chart snapshot trick is slick. I haven't tried it with Arize specifically, but I've done something similar by having a Lambda grab a quick screenshot from our monitoring dashboard. You're right about the latency, though it can be worth it for those complex drift alerts where a picture really is worth a thousand log lines.
On cold starts, we just accept the extra few seconds for alerts. For us, a model performance alert isn't usually a "server's on fire" level of urgency where seconds count. We've found keeping the Lambda warm with a scheduled ping is enough to prevent it most of the time, honestly. The bigger risk for us is the Arize API or Slack having a hiccup during that fetch.
don't spam bro
Spot on about the direct link to the investigation panel. That's the lifeline for whoever's on call. One thing I'd add is to make sure your Lambda also checks for and handles the alert's `severity` from Arize if you have it set. Mapping that to PagerDuty's urgency can prevent a low-severity data drift from waking someone up at 3 AM unnecessarily.
We also found that building a small retry with exponential backoff into the Lambda for the PagerDuty API call is worth the extra few lines of code. Their API is solid, but network blips happen.
Sleep is for the weak