Skip to content
Just built a poor m...
 
Notifications
Clear all

Just built a poor man's AI triage by piping Slack alerts to a custom GPT. Surprisingly not terrible.

18 Posts
18 Users
0 Reactions
58 Views
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
Topic starter   [#23615]

I've been experimenting with ways to introduce lightweight, intelligent triage into our alert flow without committing to a full vendor platform. Our team's primary notification channel is Slack, so I decided to see if I could route those alerts through a structured prompt to a custom GPT (using the API) for initial assessment.

The basic architecture is simple:
* A Slack webhook listener (a small Python Flask app on a free tier cloud instance) captures specific alert messages from our SIEM integration.
* It formats the alert text, along with some static context about our environment (e.g., "prioritize alerts from the finance VPC"), into a predefined prompt.
* This prompt is sent to a custom GPT configured with guidelines for security triage—asking it to classify urgency, suggest possible false positive indicators, and recommend initial containment steps.
* The GPT's response is then posted back to a dedicated Slack channel, prefixed with `[AI_Triage]`.

The results have been surprisingly coherent. For example, it correctly flagged a series of failed login alerts from a non-existent user as a likely scan (low urgency) and suggested checking the source IP against our threat intel feed, while a `CloudTrail` `StopLogging` alarm was immediately escalated as critical. It's not perfect—it sometimes hallucinates irrelevant details from its training—but as a first-pass filter to reduce alert fatigue, it's proving useful.

I'm curious if others are pursuing similar "poor man's" AI SOC integrations. Specifically:
* What prompt engineering techniques have you found most effective for security alerts?
* How are you handling the inherent risks of data leakage when sending potentially sensitive alerts to an external LLM API?
* Has anyone built a simple feedback loop to improve these models, perhaps by logging analyst overrides?

The cost is negligible so far, akin to a few reserved instances. The biggest hurdle is trust, not technology.

—A


Every dollar counts.


   
Quote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

The coherence you're seeing aligns with my own experiments using LLMs for initial signal classification. The static context about prioritizing finance VPC alerts is a smart inclusion-it mitigates the biggest weakness, which is the model's lack of live environmental awareness.

Have you measured any latency from alert generation to the triage response appearing in Slack? I tried a similar pipeline and found the API call was the bottleneck, sometimes adding 8-12 seconds. That's fine for low-urgency triage but becomes problematic if you ever want to escalate its role.

Also, how are you handling the risk of the model getting "creative" with containment steps? I had to build a strict validation layer that cross-references recommended actions against a known, approved list because mine once suggested isolating a server by shutting down a core network service we don't even operate.


Data > opinions


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Cool experiment for a weekend project. But you built a critical path dependency on a free tier instance and an opaque API.

What happens when your Flask app hits a memory leak? Or OpenAI's API has an outage during a real incident? You've now hidden your primary alert channel behind two new single points of failure.

Static context isn't enough. The model has no real-time state. It can't know if finance VPC is currently under maintenance, which makes its "prioritization" guesswork.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@alexb)
Reputable Member
Joined: 3 months ago
Posts: 257
 

You're right about the free tier and opaque API being single points of failure - that's definitely the "poor man's" part. For a weekend experiment, it's a fun way to explore the concept, but I wouldn't push it to production without redundancy.

Your point about real-time state is the bigger issue though. Even with perfect uptime, the model's context is stale the moment it's loaded. If finance VPC is under maintenance, the triage logic is backwards. You'd need a live API hook to your CMDB or maintenance calendar to make the prioritization useful, and now you're building a whole integration layer.


Data > opinions


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

So your weekend project successfully identified a failed login scan from a non-existent user. Congrats. A regex could have done that without calling an external API and adding 12 seconds of latency.

You're just adding a fancy, expensive, and unreliable filter before a human even sees the alert.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

That's a fair critique. The regex point is solid for that specific, simple pattern.

But the broader value of the experiment, to me, was seeing if the model could handle the messy, non-standard alerts where writing and maintaining a regex for every possible variation becomes a burden itself. It's less about replacing simple rules and more about assisting with the ambiguous cases that slip through them.

The latency and cost are real trade-offs, though. It makes you think hard about which alerts are actually worth that extra processing.


Stay factual, stay helpful.


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

It's clever, I'll give you that. But you're training your team to trust an API that can and will confidently hallucinate during a real crisis. I once saw a similar setup, fed an obscure database error, recommend a server restart. The actual cause was a corrupted schema migration. The "initial containment step" would have made the outage worse.

Static context about the finance VPC is worse than useless, it's dangerous when stale. That model has zero idea if that VPC is currently in the middle of a scheduled, but delayed, penetration test.



   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

You're absolutely right about the hallucination risk - trusting a black box with containment steps is a recipe for disaster. I've found you can mitigate that a bit by strictly limiting the model's output format. Instead of freeform suggestions, force it to only output a structured severity score and a classification tag from a predefined list you control.

But yeah, the stale context problem is a killer. Even with a live API feed to a CMDB, you'd need to constantly verify that the model is actually using that new data correctly. It's a fascinating prototype, but that leap from prototype to reliable system is a huge one.


Always testing.


   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

Agreed on the hallucination risk, it's the biggest blocker. Limiting output to a strict schema helps, but you can't eliminate it.

The real danger is that false confidence. The team starts leaning on a suggestion that *sounds* authoritative, especially when tired during an incident. I've seen similar setups create a "shadow playbook" that people follow without the critical thinking a human responder would have.

Your database error example is perfect. The model doesn't know what it doesn't know, so it picks a common, plausible fix. You need a human to ask why the error occurred in the first place.


Build once, deploy everywhere


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

That "shadow playbook" effect you mentioned is a real concern, and it's a people problem as much as a tech one. If the AI's suggestion becomes the path of least resistance, it can quietly erode institutional knowledge over time. The team might stop updating the actual runbooks because the model seems to handle it.

You can try to guardrail the output with schemas, but you can't guardrail human psychology during a 3 a.m. page. The goal should be keeping the human actively in the loop for diagnosis, not just using the model to generate a plausible-sounding first step.


~Harry


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Exactly. The problem isn't the wrong answer, it's the erosion of the "why." If the model suggests a restart, the team stops asking what caused the failure in the first place. Over time, you lose the institutional memory of past incidents and their root causes.

Guardrails can't fix that. You need to structure the system to force human engagement. Make the model output only diagnostic questions or links to relevant runbook sections, never action items. Keep the thinking on the human.


Metrics don't lie.


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

This is such a crucial point that often gets overlooked in the rush to automate. You've hit on the exact reason why so many "assistants" end up degrading the system they're meant to help.

Forcing the output to be diagnostic questions or runbook links is a smart design. It treats the model less as an oracle and more as a librarian for your institutional knowledge, pointing you to the right shelf but never reading the book for you. It actively pulls the human toward the documented process instead of around it.

The real test of any triage system, AI or not, is whether it makes the team's collective understanding stronger or weaker over time. If it shortcuts the "why," it's actively harmful, no matter how accurate its suggestions appear in the moment.


Stay curious.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

Nailed it. The librarian analogy is perfect. It shifts the design goal from "give an answer" to "activate the right human process."

One nuance I've seen: even linking to runbooks can backfire if the knowledge is stale. If the model points you at "Procedure for API Gateway 5xx errors" but you silently migrated to Envoy last quarter, you're still chasing ghosts. The librarian has to know which books have been checked out and not returned.

So the harder, more valuable problem becomes keeping that "library" of links and questions alive and current. The AI's real job might be flagging when its own suggested resources are outdated, which ironically means building a second layer of integrity checking around it.



   
ReplyQuote
(@elliek2)
Reputable Member
Joined: 3 months ago
Posts: 355
 

That's a really good point about the stale runbook problem. It feels like you're just pushing the trust issue one step back. The model might be great at pointing to a document, but how do we know the document itself is right?

It makes me wonder, is there any existing tool or process that's good at flagging outdated knowledge? Like, something that pings the last editor after six months to confirm it's still valid? Seems like we need that for humans too, not just the AI librarian.



   
ReplyQuote
(@henryp)
Reputable Member
Joined: 2 months ago
Posts: 294
 

Tools exist. People ignore them. We had runbooks with automatic "last reviewed" dates that triggered nag emails. The docs still rotted.

You're right that we need it for humans, too. But that's the whole problem: you've just invented a second, equally fallible system to watch the first one. At some point the audit log becomes the thing you need to audit.

What if the only reliable signal is when a procedure fails during an incident?


Doubt everything


   
ReplyQuote
Page 1 / 2