Thank you for outlining those categories. That's a very useful framework for evaluating these kinds of features. The "operational noise" category is the one that truly determines the workload, and your split seems to align with what others are seeing.
Your observation about the integration being "primarily observational" is key. It forces you into a reactive posture. Without a programmatic way to define those alert categories upfront, you can't translate your operational policies into the system. You're left chasing the noise it creates, which undermines the proactive stance the feature is supposed to enable.
That breakdown into three categories is super helpful. You mentioned two being high-fidelity and operational noise... I'm really curious what the third category was. In my experience, there's often a middle bucket - things that *look* like noise but hint at actual process problems, like unapproved SaaS tools popping up.
Also, that two-week baseline period is interesting. Did you notice if it required a "clean" period, or did it try to learn from *all* traffic, including known anomalies?
Webhooks or bust.
You're calling those data exfiltration cases "high-fidelity," but let's be precise. You validated them with external signals: destination reputation and HR context. The only new variable from this module was the volume trigger, which is just basic stat work.
So the "AI" part didn't find the threat, it just flagged an uptick in bytes. The correlation that made it actionable came from elsewhere. That means you're paying for an alert aggregator, not a detector. The noise you cut off describing is the system trying, and failing, to do the part they actually advertised.
Data skeptic, not a data cynic.