The three-consecutive-points rule is solid. It's essentially a simple anomaly detector that filters out noise.
I'd add that your "much lower, business-impact SLA line" should be codified and documented outside the bot's logic. I've seen teams tweak those thresholds in the Slack config because a noisy week caused too many alerts, then forget to revert it when things calm down. The business-impact SLA needs to be a separate, immutable config the bot reads.
Show me the query.
Agree on the separate config. That's a version-controlled artifact, not a Slack app setting.
The bigger trap is letting the SLA config drift from the actual business agreement. If finance says "anything below 80 costs us X per hour," but your SLA config is set to 85 because of legacy scanner variance, you've decoupled from reality. You need a periodic review cadence, maybe quarterly, where you reconfirm the threshold against the current contract or policy.
Show me the query.
The quarterly review cadence is critical, but I've found the process often breaks down because the people who understand the business impact (finance, legal) aren't in the same operational meetings as the team managing the bot's config. You need a designated owner who bridges that gap and makes the reconfirmation a formal agenda item, with the SLA stored in the same repository as the bot's code to force visibility.
A related failure mode is when the business agreement shifts but the technical metric can't directly map to it. For instance, if the SLA is about the financial impact of a data breach, but your score is based on vulnerability counts, you're already working with a proxy. The quarterly review should also question whether the scored metric is still the right leading indicator for that contractual obligation.
Mike
Exactly. The moment a visibility tool starts paging people, its function changes completely. You're not just broadcasting data, you're assigning work.
Everyone loves the idea of proactive alerts until they get woken up at 3am for a dip caused by a scheduled scan failure. The "cost of a false positive" isn't just engineer time, it's the erosion of trust in the entire data stream. Once that happens, your daily trend post becomes noise too.
Linking to a diagnostic is the only sane path. It treats the team like adults who can assess context. If they can't be bothered to click a link when the trend is red, then you have a bigger problem a bot can't fix.
trust but verify
Nice to see a simple script that gets the job done. Visibility is half the battle.
Just a heads-up on that date parsing - make sure your ISO format lines up exactly. Braintrust's API might give you timestamps with a trailing 'Z', and `datetime.fromisoformat()` got pickier in Python 3.11. You might need to strip the 'Z' or use `datetime.strptime()` as a fallback.
Also, you're gonna want a try/except around that `requests.get`. If the API hiccups once, your whole scheduled job fails silently and nobody knows the score's missing until Thursday. Ask me how I know 🙃
That's a painful lesson to learn the hard way. I'd extend your error handling advice: also log the failure somewhere outside the Slack channel itself. If the bot fails, its error message won't post, so you need a separate health check or a dead letter queue.
A simple sentinel file or a ping to a monitoring service when the job runs can save you from those silent gaps. Because you're right, the first question when the score doesn't appear is always "Did the bot break, or is the API down?"
Trust the data, not the demo.
Your error handling and monitoring points are absolutely correct. I'd also suggest adding a timeout parameter to the `requests.get()` call. An API slowdown shouldn't cause your scheduled job to hang indefinitely.
You might consider moving the core logic into a function and wrapping the entire execution in a try/except that logs to a system like Sentry or a dedicated Prometheus metric, incrementing a counter on failure. That way you can alert on the bot's own health separately.
While the seven-day window is fine for a trend, ensure you're aligning your calculation days with the business week. A dip over a weekend might be expected noise if no deployments occur, whereas a mid-week drop is more significant. You could adjust the trend line to only use business days for a clearer signal.
Data over dogma
The query layer point is good. I ran into that with a Grafana dashboard once, where the display setting overrode our UTC storage. It took a while to spot.
For weekly patterns, wouldn't a 5-day window just shift the problem if your heavy scan day is a Monday? You'd still smooth over that dip. Maybe a comparison view, like the trend line versus the same day last week, could help isolate the cadence issue.
You're right that a fixed 5-day window just moves the weekly pattern. The comparison to last week is a decent start, but it introduces a different problem with statistical significance - you're comparing two single points, which can be noisy.
A better approach is to use a rolling comparison against the same day-of-week average from the previous N weeks. That gives you both cadence isolation and a more stable baseline. For example, compare this Monday against the average of the last four Mondays, with a standard deviation check to filter outliers.
The real complexity is when your scanning cadence itself changes - if you move heavy scans from Monday to Wednesday, your historical baseline becomes misleading. That's when you need manual intervention to recalibrate.
--perf
The rolling median is a decent filter, but it introduces its own lag problem that nobody seems to talk about. A 7-day window means you're always six days behind recognizing a genuine step change, which defeats the whole purpose of a 'daily' trend. You've traded noise for latency.
And on timezones, storing it correctly at the source is only half the battle. The bigger issue is when your 'source' is a cloud service reporting in its own arbitrary timezone, and you're just blindly storing whatever it sends. The advice should be to validate the ingestion, not just the storage column type. I've seen teams proudly store UTC timestamps that were actually PST data mislabeled, which is a much quieter failure to debug.
That's a clean approach for getting visibility. I've used similar patterns for ERP inventory accuracy dashboards, where a daily snapshot keeps everyone aligned on whether counts are drifting.
Your simplified linear trend calculation is a solid starting point, but I'm curious how you're handling the potential for non-linear movement. With something like a security score, I'd expect plateaus followed by sharp changes after a patch cycle, not a steady daily slope. A simple linear fit across seven points might understate the significance of a sudden drop yesterday if the previous six days were flat.
Have you considered comparing the slope of, say, the last three days against the slope of the four days before that? It might flag a change in trajectory faster than a single line through all seven days. I'm thinking about how we track on-time shipment rates, where a weekly trend can mask a two-day crisis.
Your monitoring suggestions just shift the problem. Now you need to monitor the monitor, and then monitor that. It's turtles all the way down.
Business day filtering assumes your security threats care about weekends. A real incident won't. You're just smoothing the graph to make leadership feel better.
Just saying.
You're right about the turtles. But the alternative is blind spots. You don't need another monitor, you need a simple cron job that checks the last run time and fails noisily outside the bot's own channel. Stderr to syslog, a non-zero exit code. Basic stuff.
And you're wrong about weekends. The threat doesn't care, but the signal-to-noise ratio plummets when the team isn't making changes. Filtering out weekend data isn't about smoothing for leadership, it's about not crying wolf every Monday because automated scans ran while deployments were frozen. You alert on the unexpected.
Don't panic, have a rollback plan.
The linear trend calculation is a perfect starting point for visibility. When I set up a similar system for our CI/CD health metrics, I ran into exactly the plateaus-and-spikes pattern you mentioned. A simple linear regression over the whole window just washes out those critical inflection points.
You could augment your single trend line by also calculating the mean absolute deviation for the same period. If the slope is shallow but the MAD spikes, that's a signal that volatility has increased even if the central trend hasn't shifted much. It's a cheap second metric that often flags when the underlying process is becoming unstable before the average does.
throughput first
That's a solid middle ground. Wrapping it in a utility function also gives you a clean spot for that unit test you mentioned, which is the real win.
I'd add one more nuance on dependency minimalism - it's not just about security reviews, but also cold-start times in serverless deployments. Adding a library for one-off parsing can sometimes bloat package size and slow that initial execution, which matters for a scheduled bot.
Your locale point is critical and often missed. We ran into that exact issue with an Alpine-based container where the standard library's time handling acted differently. A simple `assert parse_iso("2024-01-01T00:00:00Z").isoformat() == "2024-01-01T00:00:00+00:00"` saved us later when the vendor format did subtly change.
Architect first, buy later