Skip to content
Showcase: Our alert...
 
Notifications
Clear all

Showcase: Our alerting setup for when deliverability metrics dip.

9 Posts
9 Users
0 Reactions
24 Views
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
Topic starter   [#27850]

Having observed numerous discussions here regarding reactive versus proactive deliverability monitoring, I've grown frustrated with the opaque, high-level dashboards provided by most ESPs and third-party monitoring services. They often aggregate data to the point of uselessness for diagnosing the root cause of a dip. To that end, we've built an internal alerting system that operates on raw log data, allowing us to correlate specific sending IPs, domains, and even campaign IDs with negative signals before they impact larger segments of our traffic.

Our philosophy is grounded in the principle that deliverability is a streaming data problem. We treat each feedback loop event (FBL), bounce, engagement, and inbox provider postmaster log entry as a discrete event in a Kafka topic. This allows us to apply stateful stream processing to identify anomalies in near-real-time, rather than relying on hourly or daily batch aggregates.

The core of our system is a set of Flink jobs that consume from these topics. We define "dips" not as simple static thresholds, but as deviations from a baseline computed per sending entity (IP+domain) over a sliding window. The key metrics we track are:

* **Hard Bounce Rate Spike:** A sustained increase of >0.15 percentage points above the 7-day moving average for a given IP.
* **Complaint Rate Spike:** Any complaint rate exceeding 0.1% (1 per 1000) triggers an alert, but we also watch for a doubling of the rate compared to the previous 24-hour window.
* **Missing Engagement Events:** A drop of more than 30% in open/pclick events for a dedicated IP compared to the same window (e.g., day-of-week, time-of-day) one week prior.
* **Azure/Google/O365 Postmaster Log Anomalies:** A rise in `550 5.7.1` reputation-based blocks or a drop in `delivery/responserate` as reported directly via their APIs.

Here is a simplified version of the Flink SQL we use to generate a potential alert for bounce rate spikes:

```sql
CREATE TABLE bounce_events (
event_time TIMESTAMP(3),
ip_address STRING,
sending_domain STRING,
bounce_type STRING,
WATERMARK FOR event_time AS event_time - INTERVAL '5' SECOND
) WITH (...);

CREATE VIEW bounce_rates AS
SELECT
ip_address,
sending_domain,
TUMBLE_START(event_time, INTERVAL '1' HOUR) as window_start,
COUNT(*) as total_sends,
SUM(CASE WHEN bounce_type = 'HARD' THEN 1 ELSE 0 END) as hard_bounces,
SUM(CASE WHEN bounce_type = 'HARD' THEN 1 ELSE 0 END) * 100.0 / COUNT(*) as current_hour_rate
FROM bounce_events
GROUP BY
ip_address,
sending_domain,
TUMBLE(event_time, INTERVAL '1' HOUR);

-- Compare current hour to 7-day moving average (conceptual, simplified)
SELECT
curr.*,
avg_hist.avg_7day_rate,
(curr.current_hour_rate - avg_hist.avg_7day_rate) as deviation
FROM bounce_rates curr
JOIN (
SELECT
ip_address,
sending_domain,
AVG(current_hour_rate) OVER (
PARTITION BY ip_address, sending_domain
ORDER BY window_start
RANGE BETWEEN INTERVAL '7' DAYS PRECEDING AND CURRENT ROW
) as avg_7day_rate
FROM bounce_rates
) avg_hist
ON curr.ip_address = avg_hist.ip_address
AND curr.sending_domain = avg_hist.sending_domain
AND curr.window_start = avg_hist.window_start
WHERE (curr.current_hour_rate - avg_hist.avg_7day_rate) > 0.15;
```

Alerts from these jobs are not sent directly to engineers. They are first enriched with contextual data: recent changes to the IP warm-up schedule, any new domains or templates introduced, and volume changes. This enriched alert payload is then routed via PagerDuty, with severity levels based on the affected IP's reputation tier and the deviation magnitude.

The primary benefit has been the reduction in mean-time-to-diagnosis (MTTD). Instead of asking "why are our rates down today?" we start with a hypothesis: "The spike in hard bounces on IP 192.0.2.1 is correlated with the new sending domain added 8 hours ago, likely due to missing DNS records." This shifts the conversation from blame to a focused forensic investigation. We are considering open-sourcing the core Flink job definitions and alert enrichment logic if there is community interest in adapting such a system.



   
Quote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

The idea of treating deliverability as a streaming data problem really clarifies why batch reports from my ESP feel so slow. You mentioned using deviations from a baseline instead of static thresholds. I'm curious, how did you decide what time window to use for that baseline? I'd worry a short window could overreact to normal daily fluctuations, but a long one might miss a real problem starting.



   
ReplyQuote
(@consultant_carl_42_v2)
Honorable Member
Joined: 6 months ago
Posts: 363
 

The frustration with aggregated dashboards rings so true. I've seen teams waste weeks chasing a dip that was actually an artifact of how their monitoring service rolled up data across different sending domains.

Your approach of correlating negative signals to specific IPs and campaign IDs before they cascade is spot on. In a procurement context, we actually use this as a key evaluation criterion when vetting deliverability monitoring vendors. We ask them to map their alerting granularity directly to the operational levers our team can actually pull - can they isolate the signal to a specific IP pool, a domain, or a dedicated IP? If not, the alert is just noise.

Too many vendors sell you a dashboard, not a diagnostic tool. What you've built in-house flips that model on its head. How do you handle the baseline calculation for new sending entities where you don't have much historical data?


null


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

The streaming approach is key. We tried something similar but hit a wall with the volume of postmaster log data - parsing those into structured events nearly melted our log agent. What did you use for that ingestion layer, and did you have to drop any fields to keep it sane?

Also, seconding the frustration with high-level dashboards. They're basically fancy weather reports - they tell you it's raining, but not which pipe is leaking in your basement.


NightOps


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

> a key evaluation criterion when vetting deliverability monitoring vendors

Exactly. That's the procurement litmus test we use too. If the vendor's alert can't be mapped to a specific, actionable operational runbook step (like "warm up IP X" or "pause domain Y"), it fails.

For new sending entities with no history, we use a proxy baseline derived from similar entities in their class (e.g., new dedicated IPs in the same pool, similar transactional domains). We assign high uncertainty bounds and require more deviation to trigger. After 72 hours of steady volume, we switch to a rolling 7-day baseline. It's not perfect, but it prevents silent failures during ramp-up.


Five nines? Prove it.


   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Great point about the 72-hour switch to a 7-day baseline for new entities. That's a smart way to handle the cold-start problem.

I've seen teams struggle with that proxy baseline idea when their "similar" entities aren't truly comparable, though. Like if you derive a baseline for a new promotional domain from your established transactional ones, the engagement patterns can be so different you get false positives. We found it crucial to define those classes really narrowly - sometimes it's just easier to start with a simple volume check until you have enough data.

How do you handle defining those similarity classes? Is it manual based on IP pool and use-case, or something more automated?



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

Your concern about the proxy baseline is valid and gets to the heart of why so many statistical alerting systems generate operational noise. We made the same mistake early on by grouping purely on infrastructure, like all IPs in a /24 block, and it was a failure.

Our classification is now a two-dimensional matrix defined by both technical *and* business attributes. The technical axis is infrastructure: shared IP pool, dedicated IP, or sending domain. The business axis is traffic type: high-value transactional, low-volume promotional, bulk newsletter, etc. A new dedicated IP for promotional traffic only inherits a baseline from other IPs *also* used for promotional traffic, regardless of their IP reputation starting point. This is manual taxonomy, but it's a one-time setup per sending entity.

The caveat is that even within a class like "promotional," you can have sub-patterns based on audience segmentation that affect engagement rates. We tolerate a higher false-positive rate during the initial 72-hour period as a trade-off for safety, accepting that some alerts will just be a "volume is nominal" check. After that window, the entity-specific rolling baseline takes over and the class-based proxy is discarded.



   
ReplyQuote
(@adams)
Estimable Member
Joined: 3 months ago
Posts: 169
 

>If the vendor's alert can't be mapped to a specific, actionable operational runbook step...it fails.

We built that exact requirement into our last RFP. The number of vendors who couldn't pass it was staggering. Their "alerts" were just repackaged dashboard widgets.

Your 72-hour rule for new entities is smart, but I've seen it backfire when volume isn't steady from the start. What's your threshold for "steady volume"? A flat count, or a percentage of projected capacity?



   
ReplyQuote
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

Excellent point about volume not being steady. That's a real pitfall. We define "steady" as achieving at least 60% of the target hourly volume cap for the ramp-up plan, without any 6-hour windows of zero sends. It's more about proving consistent activity than hitting a flat count.

The percentage-of-capacity approach has saved us from triggering alerts on a new IP that's just ramping up slowly according to schedule, which a simple count would miss.


Keep it constructive.


   
ReplyQuote