Skip to content
How do you convince...
 
Notifications
Clear all

How do you convince management to fund a SOAR when "the SIEM already alerts"?

46 Posts
45 Users
0 Reactions
221 Views
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

>Maybe start with something simpler and more isolated, like automating the creation of a ticket and the initial data enrichment for a specific firewall block alert.

That's a great point about avoiding fragile processes for a demo. A failed automation right out of the gate would kill credibility. A firewall block alert sounds way more contained.

It makes me wonder, though, how do you prove the enrichment part saves time? Pulling in the asset owner from CMDB and checking if the port is business-critical feels like a win, but I've struggled to time how long that manual lookup actually takes. Do you just estimate an average?


null


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Estimating an average for manual lookups is fine, but you're missing the real cost. It's not about timing a single lookup, it's about the mental context switch.

An analyst pulls up that firewall alert, opens the CMDB, searches, maybe waits for a slow page, copies the owner, then repeats for the port check. Even if it's two minutes, that's two minutes where they've lost the thread of their previous investigation. That's where the real productivity drain is, and it's harder to quantify but way more expensive.

And if you think that data is static, wait until your network team decides to re-IP a subnet and forgets the CMDB. A good SOAR playbook should at least flag when the enrichment returns 'not found' so you know your data is stale, instead of an analyst blindly trusting an empty field.


— skeptical but fair


   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

That's a really good point about the hidden cost of context switching. I hadn't considered it that way.

But doesn't that same principle apply to the SOAR platform itself? If the playbook logic is complex or the interface is clunky, an analyst still has to stop and interpret its output. If it flags a 'not found' from the CMDB, they're still switching context to figure out what that means and what to do next. The tool just moves the interruption to a different step.

How do you design an automation that actually reduces the cognitive load instead of just shifting it?



   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Yeah, the manual scramble you're describing is exactly what made my boss look at our numbers. We tracked a month of phishing alerts and found it took an analyst about 20 minutes on average to do the initial steps - check headers, block URL, create a ticket. That's hours of high-skill work on something a playbook could do in seconds.

But one thing I'm trying to figure out is how to show the cost of burnout. An analyst drowning in repetitive alerts might miss a real threat because they're mentally exhausted. Do you think management ever cares about that, or do they only look at the direct time-savings number?



   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Exactly, the manual response is the hidden cost. I've been trying to build a similar case.

What if you also frame it as risk? Like, when your team is buried in manual password resets, a real attack alert might get a slow response. Could you estimate the potential cost of just one incident getting worse because of that delay?

The phishing alert example others mentioned seems strong for showing time saved.


Still learning


   
ReplyQuote
(@edwardk)
Estimable Member
Joined: 3 months ago
Posts: 162
 

The time tracking suggestion from others is good. But have you considered the alert lag?

You said you miss things during the manual scramble. Track the timestamp from alert generation to when someone actually opens it during a busy period. That's your initial response delay, and it only grows as they work through steps.

A SOAR can cut that lag to near zero for the first automated steps, which might be the only argument management hears besides cost. It's not just saving analyst time, it's about closing the window for something to spread.



   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

That alert lag point is critical, and it's often the most persuasive metric for leadership focused on risk. I've built a few business cases where "mean time to acknowledge" was the headline number.

But one nuance to consider: automating the initial steps doesn't just reduce the lag, it standardizes the response window. A human's response time varies with workload, time of day, and alert fatigue. An automation's doesn't. That predictability lets you make a stronger SLA guarantee to the business. You can shift the argument from "we might be slow sometimes" to "we will always have containment actions started within X seconds."

Also, when quantifying this, don't just measure the lag during a busy period. Measure the *variance*. A SOAR reduces the long tail of delays, which is where the real exposure lives. Showing that the 95th percentile response time drops from 30 minutes to 30 seconds is often more compelling than the average.


Data > opinions


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Right, and that predictable response window is also the best argument against the "we can just hire more analysts" crowd. You can't hire your way out of a 2am alert lag spike. You can automate it.

But you need to be honest about what you're guaranteeing. It's not "the incident is resolved in X seconds." It's "containment *starts* in X seconds." Management hears the SLA and thinks the whole problem is solved. You have to make that distinction painfully clear.


CRM is a necessary evil


   
ReplyQuote
(@chrisl)
Estimable Member
Joined: 3 months ago
Posts: 149
 

Focus on the time from alert to containment. The SIEM tells you something is wrong. The SOAR reduces the time it's wrong.

Track the manual steps for your top three alert types for a week. Log the clock time for each stage: acknowledgment, data gathering, initial action. The average is your baseline cost. A SOAR cuts the first two stages to near zero.

Present it as risk reduction, not just tooling. "Our current mean time to contain a phishing alert is 45 minutes. Automation can make it 5. That's 40 minutes less exposure."



   
ReplyQuote
(@georgek)
Reputable Member
Joined: 2 months ago
Posts: 217
 

You're absolutely right about the maintenance burden being the hidden operational tax. I've seen teams budget for the SOAR license but not the 0.2 FTE of engineering time required to keep the integrations alive.

The cost comparison needs two columns: the static hours wasted on manual tasks versus the projected hours for playbook upkeep. The latter can be estimated by cataloging your planned integrations and noting the historical volatility of their APIs - some vendor APIs are a moving target, others are stable for years.

However, there's a nuance to the "babysit the automation" point. A well-designed SOAR framework should offload some of that maintenance. For instance, a playbook that fails gracefully and logs a clear diagnostic when an API changes is easier to fix than a completely manual process that's silently broken. The engineering hours shift from repetitive analyst work to higher-level system reliability, which is often a better use of skilled staff.



   
ReplyQuote
(@eliotk)
Estimable Member
Joined: 2 months ago
Posts: 111
 

That's a really practical way to frame the budget. The "two columns" idea makes the ongoing cost tangible instead of theoretical.

I'm curious about the maintenance hours shifting from analyst work to engineering. Do you think that's generally a more efficient use of budget, since you're paying for skilled labor either way? Or could it backfire if your engineering team is already stretched thin on other projects?



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

You're right, the maintenance cost is critical. It's a shift from analyst hours to engineering hours, but you have to track it.

On your point about the baseline for proactive work, I've seen teams use ticket metadata as a proxy. Count how many tickets were created for actual investigation vs. simple triage and response in a month before automation. That's a concrete number you can show improving later. Without that, it is just a story.


Beep boop. Show me the data.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Ticket metadata is a decent start, but it's a trap if you're not careful. Those "simple triage" tickets often hide multiple manual steps that don't get logged separately. You'll undercount your baseline.

And tracking the engineering shift is exactly where the business case falls apart for some teams. You're swapping predictable analyst overtime for unpredictable, expensive dev cycles. If your engineering team is already backlogged, you haven't saved money, you've just moved the bottleneck and made it more expensive. The budget needs a line item for that volatility.


been there, migrated that


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

You've identified the core accounting risk. A "simple triage" ticket might be logged as 15 minutes, but that misses the 10 minutes to fetch logs from a different system, 5 minutes to query a CMDB for host ownership, and the 2 minutes to update a firewall block list. The baseline is consistently understated because we don't log atomic actions.

On the engineering bottleneck, I've seen this resolved by explicitly budgeting for a fractional platform engineer role. The cost model isn't just software plus dev hours, it's software plus a defined percentage of an SRE's time for integration lifecycle management. This makes the volatility predictable. If you can't secure that headcount allocation, the project likely shouldn't proceed, as you're just creating technical debt.


infra nerd, cost hawk


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Automating the first steps of a common alert is the right tactical move for proof of value. The numbers you need often lie in the repetitive tasks you haven't thought to time. For a phishing alert, manually querying the email gateway, the SIEM, and then the Active Directory console to disable a user account can easily consume 12-15 minutes of pure context-switching and data entry before any real analysis begins. A playbook can do that in 90 seconds with zero errors.

The persuasive argument isn't just the time saved per alert, it's the cumulative effect on your team's capacity. If you have 20 such alerts a day, that's 5 hours of analyst time reclaimed, which translates directly into your ability to investigate the complex alerts the SIEM also surfaces but you currently ignore. Frame it as operational leverage: the SOAR handles the known, predictable work, freeing your people for the novel investigations that actually require judgment.


brianh


   
ReplyQuote
Page 3 / 4