Skip to content
Notifications
Clear all

Built a Slack bot that posts our daily security score trend

63 Posts
58 Users
0 Reactions
280 Views
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
Topic starter   [#23400]

Hey folks! 👋 I've been working on a little automation project that's been a hit with our security team, and I wanted to share the approach. We use Braintrust to track our security posture score, and I built a Slack bot that posts the daily trend line every morning. It gives everyone a quick, visible pulse on whether we're improving or if something needs attention.

The core idea is simple: use Braintrust's API to fetch the score history, calculate the recent trend, and format a clean message for Slack. Here's the key part of the Python script that runs as a scheduled job:

```python
import requests
import os
from datetime import datetime, timedelta
from slack_sdk import WebClient

BRAINTRUST_API_KEY = os.environ.get("BRAINTRUST_API_KEY")
SLACK_BOT_TOKEN = os.environ.get("SLACK_BOT_TOKEN")
PROJECT_ID = "your-project-id-here"

# Fetch last 7 days of scores
response = requests.get(
f"https://www.braintrust.dev/api/project/{PROJECT_ID}/scores",
headers={"Authorization": f"Bearer {BRAINTRUST_API_KEY}"},
params={"limit": 7}
)
scores_data = response.json()

# Simple linear trend calculation (simplified for example)
dates = [datetime.fromisoformat(s['date'][:-1]) for s in scores_data['scores']]
values = [s['value'] for s in scores_data['scores']]
if len(values) > 1:
trend = values[-1] - values[0] # change over the period
else:
trend = 0

# Build Slack message
emoji = "📈" if trend >= 0 else "📉"
client = WebClient(token=SLACK_BOT_TOKEN)
client.chat_postMessage(
channel="#security-metrics",
text=f"{emoji} *Security Score Trend* (7-day window): {trend:+.2f}nLatest score: {values[-1]:.1f}"
)
```

A couple of best practices I learned:
* **Always cache or limit API calls.** Braintrust's API is responsive, but be nice to it. I fetch only the last 7 data points.
* **Handle missing data gracefully.** Some days might not have a score entry, so my production code has logic to interpolate or skip gaps.
* **Make the Slack message actionable.** We added a direct link to the Braintrust project dashboard in the message using `blocks` for richer formatting.

The biggest "aha" was deciding on the trend window. A 7-day rolling window smoothed out weekend noise but still showed recent shifts. We also experimented with posting to different channels for devs vs. leadership, with slightly different commentary.

It's a small thing, but it's made our security metrics way more visible and sparked some great conversations. Has anyone else built similar integrations? I'm curious about how you're calculating trends or if you're pulling in other Braintrust data.

Happy coding!


Clean code, happy life


   
Quote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

Watch your runtime costs if you're polling that API daily with a scheduled job. Consider moving the logic to a serverless function that only runs when a new score is actually available, maybe via a webhook from Braintrust. You'll cut compute time by 99% and avoid paying for idle time.


Show me the bill


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

The cron job's fine. That's maybe $1 a month if you're overspending. Now you've got to manage webhook registration, retry logic, and another failure domain. Is cutting a dollar's worth of compute really worth the extra complexity?


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Nice! I've been thinking about building something similar for our vulnerability metrics. I like how you're calculating the trend client-side with Python. One thing I'd check: does Braintrust have a dedicated endpoint for trend data? Sometimes platforms expose that directly so you don't have to pull the raw history and calculate yourself.

Also, have you thought about formatting the Slack message as a simple line chart using their block kit? You could use the `mrkdwn` fields with some creative text characters to visualize the trend inline. It's a bit more engaging than just posting the numbers.


Webhooks or bust.


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 3 months ago
Posts: 421
 

Hey, thanks for sharing the actual code snippet! Seeing the specific API call helps a lot.

> does Braintrust have a dedicated endpoint for trend data?
That's a really good question to ask. I've found that a lot of vendors *do* calculate this server-side but don't expose it, so you end up paying for the data transfer and compute twice. A quick check of their API docs could save you some future headaches if they add one.

I'm curious, how are you handling weekends or days when a new score isn't generated? Does your script just skip posting, or does it interpolate? That's where I've seen a lot of these dashboard bots fall down.


Trust the data, not the demo.


   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

Your trend calculation is going to break if your `scores_data` isn't sorted chronologically. That API response *probably* comes back newest-first, but you can't rely on "probably" with these vendor APIs. Sort your `dates` and corresponding scores explicitly before running any linear regression, or you'll post nonsense slopes one day.

Also, that `datetime.fromisoformat(s['date'][:-1])` hack to strip the 'Z' is brittle as hell. Better to handle the possible timezone suffix with a proper parser or at least a try/except block.


It's just pattern matching


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That's a smart approach, and making the security score visible like that is great for team awareness. I'm glad you mentioned it's been a hit.

One small consideration on sharing the code snippet: you'll want to double-check that it's safe to post your script structure exactly, especially if it shows the API endpoint format and any environment variable names. Sometimes even that can give bots hints for probing, or you might have company policies about sharing internal tool structures. It's probably fine, but it's a good habit to do a quick review.



   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

The point about company policies is often overlooked. People focus on technical secrets like keys, but the structure of an internal automation can itself be a blueprint for social engineering. Someone could use that script layout to craft a convincing phishing email to your ops team.

You're right that it's probably fine here. The real risk is when scripts reveal internal naming conventions, directory structures, or approval workflows.


Trust but verify — especially the fine print.


   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Hey, this is a neat project. I've built a few similar notification bots for different metrics.

One thing I'd add about your code snippet: you might want to handle the case where the `response.json()` call fails. If the API returns an error or an unexpected content type, it'll throw an exception and your whole script crashes. Wrapping that in a try/except and maybe logging the raw response on failure will save you some debugging headaches when it inevitably happens at 3 AM. 😅

Also, for the linear trend calculation, are you using a library or rolling your own? I found that using `numpy.polyfit` for a quick linear regression is more reliable than trying to implement it manually.


Integration Ian


   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Good catch on the vendor calculation point - it's frustrating how often that happens. I've seen teams burn cycles rebuilding metrics that already exist just because they're not in the primary API docs.

> how are you handling weekends or days when a new score isn't generated?
That's the real rub, isn't it? If you skip posting, the team might think the bot broke. If you interpolate, you're presenting synthetic data as fact. My rule of thumb is to post the last known score with a clear indicator like "(as of Friday)" in the message, but never fabricate a new data point. It keeps the rhythm without misleading anyone.

Have you found a better pattern for handling gaps?



   
ReplyQuote
(@carols)
Estimable Member
Joined: 2 months ago
Posts: 142
 

Moving to serverless for webhooks often shifts costs rather than eliminates them. You're trading compute time for integration complexity, which has its own operational overhead. The break-even point on that investment depends heavily on how stable the vendor's webhook system is. If their delivery is flaky, you'll spend more on monitoring and retry logic than you'd ever save on compute.


Buy once, cry once.


   
ReplyQuote
(@chrisl)
Estimable Member
Joined: 3 months ago
Posts: 149
 

Agree on the cost shift. We've seen that exact scenario where moving a daily poll to a vendor webhook created more alert fatigue than it saved. The webhook's occasional 5XX errors meant we had to build a side channel to verify delivery, which defeated the purpose.

Reliability becomes a hidden tax. If your vendor's SLA for webhook delivery is lower than your polling script's uptime, you've added risk.



   
ReplyQuote
(@connork)
Reputable Member
Joined: 3 months ago
Posts: 216
 

Yeah that hidden tax is real. It's like you solve one problem but create two more monitoring jobs.

So what did you end up doing? Go back to polling, or build a whole reconciliation system?



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

We compromised. We kept the scheduled poll but added a lightweight webhook listener as a secondary source. The poll runs daily at a consistent time. If the webhook delivers an event first, we use that data and skip the next scheduled poll for that day. It's a simple cache with a 24-hour TTL.

This gives us the timeliness of webhooks when they work, without the alert storms when they fail. The polling remains as our fallback source of truth. It's not elegant, but it works. The cost is slightly higher compute time for the polling function, but we avoid building a complex reconciliation system.

You still need to monitor both channels, but the webhook failures become informational rather than critical alerts.


Less spend, more headroom.


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Love the hybrid approach. It's the classic "event-driven when possible, scheduled as backup" pattern that works so well for this kind of reporting.

Your point about webhook failures becoming informational is key. That mental shift from "everything is broken!" to "we already have the data" saves so much ops stress. It turns a reliability problem into a minor annoyance.

I'd just add one watch-out: make sure the cache invalidation logic is rock solid. I once saw a similar setup where a missed webhook and a failed poll on the same day left a stale score sitting for 48 hours. A simple last-updated timestamp check in the polling job can catch that.


measure twice, ship once


   
ReplyQuote
Page 1 / 5