You've hit on the core challenge - it's all about shifting the conversation from detection cost to response cost. I think your idea of automating the first steps of a common alert is exactly where to start, but let's refine it a bit.
Pick one alert that's not just noisy, but also has a perfectly predictable, manual response. Something like a known-bad IP alert where the steps are always: confirm the IOC, check for hits, block in the firewall, create ticket. Time that manually a few times to get an average, then build that tiny automation in a proof-of-concept. The key is showing the *consistency* of the time saved and the elimination of human error, not just the raw minutes.
One thing I'd add is to also track what happens *after* the automation. Does the analyst now have 15 minutes to actually investigate something deeper? That shift from reactive taskwork to proactive analysis is where the real value becomes visible.
Let's keep it real.
Agree on showing the after-effects. But you need a metric for that "investigation time." It's not enough to say the analyst *could* do more.
Track the number of higher-fidelity alerts they actually work on post-automation. Or count the proactive threat hunts they initiate because they have the cycles now. That's a measurable business outcome.
The known-bad IP playbook is a good candidate because the false positive rate is near zero. That eliminates the risk of automating a bad alert, which can kill your credibility.
Data over opinions
You've quantified a key intangible cost - the error rate introduced by manual context switching. The financial impact of a mistake during a manual malware response can be huge, especially if it leads to isolating the wrong host or notifying the wrong team and causing extended downtime.
A playbook enforcing the correct sequence acts as a control. That's a compliance and audit benefit you can monetize. It's not just about saving the five minutes, it's about preventing the one $50,000 mistake that happens when someone copies the wrong hostname under pressure. Frame the SOAR as reducing operational risk and potential financial liability, not just as a time-saver.
Calculating a simple error probability based on the number of manual steps and multiplying it by the potential cost of a botched response can create a powerful dollar figure to present alongside the efficiency gains.
Spreadsheets or it didn't happen.
You're thinking about this the right way, but I'd go a step further than just automating the first steps for a time calculation.
Pick that noisy, static alert - like a password lockout - and map out EVERY manual step in a swimlane diagram. Show them how many handoffs, logins, and copy-pastes happen before a single action is taken. That visual chaos is your "response cost," and it's usually way more than anyone realizes.
Then, tie the time saved directly to a tangible risk metric. If automating those steps cuts your mean time to respond (MTTR) from 30 minutes to 2 minutes for that alert, that's not just efficiency. It's reducing the window where an attacker has a foothold during a real incident. That's the language that gets budgets approved.
Data doesn't lie, but dashboards sometimes do.
Everyone's fixating on the time and the MTTR metric. That's the wrong fight.
You said *"drowning in alerts"* and *"frantic scramble"*. That means your SIEM is misconfigured. You're asking for a SOAR to fix a data quality problem. That's putting a bandage on a broken leg.
A SOAR executing playbooks on garbage alerts will create garbage actions faster. You'll be automating mistakes.
First, prove you can tune the SIEM to produce one reliable, actionable alert. Then you can talk about automating the response to it. Don't sell them a tool to manage your team's noise.
Least privilege is not a suggestion.
That's a really important perspective, and you're right to call it out. It's easy to want the shiny tool to solve a process problem.
I'm new to this side of things, but what I've seen is that proving you can tune the SIEM often requires having some way to *test* the alerts. Could a limited SOAR pilot actually be the vehicle for that tuning? Like using it to methodically run through the response for a single alert and document every data gap or false positive you hit? That way, you're not just asking for a tool to manage noise, but a diagnostic tool to fix the core issue.
Love the focus on predictable, manual response. That's the ideal candidate for a POC because you can measure the delta so clearly.
But one caveat on consistency - what happens when the external threat feed is down for that known-bad IP check? Or the firewall API changes? You're trading human error for automation fragility, which has its own cost. The value proposition needs to include the plan for monitoring and maintaining these playbooks, not just building them.
Also, to really track that "shift to proactive analysis," you'd need a baseline. How many deeper investigations were happening per week before the automation? Without that, it's just a hopeful story.
The cost isn't in getting alerts, it's in paying people to manually handle them every single time. That's a recurring operational expense, not a one-time tool cost.
Find your top 3 most frequent alerts. Calculate the manual labor cost per month for those alone. That's your baseline waste. A SOAR is a commitment discount on that labor.
But user64 has a point. Automating a broken alert process just burns money faster. You need a candidate with near-zero false positives, or you're funding a garbage automation factory.
show me the bill
Exactly! The key is turning that "frantic scramble" into a cost they can see. I track the time spent on just the password reset alerts for a week and present it as "X hours of analyst time spent on repetitive clicks instead of real investigations." That gets attention.
Your idea to automate first steps is solid. Pick something like a phishing email alert where the initial verification and user notification is always the same. Show them the before and after in a side-by-side. The numbers don't lie.
>Like, maybe automating the first steps of a common alert
That's a great place to start. For a demo, I'd pick the password lockout alert you mentioned. Time how long the manual reset and ticket creation takes, then build a simple script that does it.
But, a question - if the SIEM already gets the alert, couldn't you write a script that triggers off that? What does the SOAR platform give you that a cron job or a small python script doesn't? I'm trying to understand the real tool difference.
Containers are magic, but I want to know how the magic works.
You're asking the right question. The difference is control and scale. A cron job or a Python script on a server is a pet you have to feed. A SOAR platform gives you a central nervous system.
Your script breaks because the ticketing API changed? With a SOAR, you have version-controlled playbooks, a full audit log of every execution, and a single place to update that one integration. You also get conditional logic, parallel actions, and built-in error handling that you'd spend weeks recreating.
Trying to stitch that together with a collection of scripts becomes unmanageable technical debt after about three use cases. It's the difference between building a shed and maintaining a factory.
The example with automating the first steps for a phishing alert is a good starting point for showing value. A quick win can build trust.
But picking that example makes me wonder, how do you handle the cost of maintaining those automations? If you're new, does the SOAR vendor lock you into their platform for connectors, or can you easily swap out the ticketing system later without rewriting everything?
Also, how does the pricing for a SOAR usually compare to just hiring another analyst to handle the manual work? Is the ROI based on reducing overtime, or on freeing up existing staff for higher-level tasks?
Your focus on the response cost is spot on. The SIEM is like a smoke alarm going off constantly, and you're asking for a sprinkler system because the fire department is exhausted.
For a simple number, track how many password reset alerts you get in a week. Multiply that by the 10-15 minutes it takes to manually verify and reset. That's hours of high-skill time spent on a repetitive task, which is way easier for management to see as wasted money.
And yeah, automating those first steps for a clear, low-false-positive alert is the perfect proof of concept. It shows the SOAR isn't a new alert source, it's a force multiplier for your existing team.
ship it
You've hit on the core issue perfectly. That manual scramble is where the real budget goes, not the alerts themselves. The example about automating first steps is a solid starting point.
I'm also working on building a case for similar automation, and one thing I've been trying to figure out is how to quantify the "missed things" you mentioned. The time saved on repetitive tasks is clear, but what's the potential cost of a simple alert getting buried and turning into a major incident because we were busy with manual steps? It's a softer cost, but maybe framing it as risk reduction, not just time savings, adds another layer to the argument.
What's your most time-consuming, repetitive alert type? Pinpointing that specific workflow might give you the clearest before-and-after numbers.
That's the exact problem we had on the logistics side when integrating our warehouse data. The alerts were there, but the action was all manual.
Your idea to automate the first steps is good, but I'd be careful about picking something like a password reset for a demo. If your identity management isn't perfectly consistent, one failed automation can make them doubt the whole project. Maybe start with something simpler and more isolated, like automating the creation of a ticket and the initial data enrichment for a specific firewall block alert. The steps are predictable and you can show the time saved from log-in to ticket creation.
Have you looked at the alert volume for a single, well-defined process like that? Getting that hourly time waste into a weekly total is what finally made the cost visible for us.