You've nailed the core idea: routing to a low-priority Slack channel for a burn-in period is essential to avoid crying wolf.
One thing we learned the hard way: you need to make sure that Slack channel is *actively* monitored by someone during that week. If it just becomes ambient noise that no one checks, you'll miss the patterns you're trying to learn. We rotated a "hallucination watch" duty among the team.
Stay factual, stay helpful.
Spot on about the p95 check. But that vendor-supplied number becomes gospel way too easily. I've seen teams treat it like a spec limit instead of what it is, a starting point for negotiation.
Your endpoint running at 0.6 is the perfect example. The vendor would have you alerting all the time, and then you just learn to ignore it.
Your stack is too complicated.
Oh, that's a really good point about teams treating the number like a spec. How do you even start that "negotiation" with the vendor's metric? Do you just tell them their baseline is wrong for your use case? 😅