Hey everyone, I'm pretty new to the SOAR side of things. I've been setting up some basic Cortex XSOAR content packs at work, but I'm struggling to visualize what a *real*, useful playbook looks like in practice beyond the simple "fetch and alert" demos.
Could someone share a concrete example of a workflow you've built or use? I'm especially curious about:
- How you handle a common alert, like a phishing email or a suspicious login, from ingestion to closure.
- What integrations you use most (like ticketing, DNS, Active Directory).
- Any logic or decisions you've added that saved real time.
Just trying to bridge the gap between the tutorials and actual daily use. Thanks in advance for any insights! 🙏
Absolutely, that gap between tutorials and production is a real thing. Let me walk you through a phishing email playbook we've been running for about a year now. The core concept is enrichment and pre-qualification to keep the vast majority of alerts from ever needing a human.
Our workflow triggers from our email security gateway alert. The playbook's first job is enrichment: it pings VirusTotal for the URL hash, checks the sender domain against our internal domain allow-list, and runs the attached file hash (if any) against our EDR's global threat intelligence. We set conditional checks after each step. If the VT score is above a certain threshold *and* the domain isn't internal, it auto-quarantines the message via O365, adds the IOCs to our block lists in the firewall and DNS, and creates a ticket in Jira Service Management with all the context pre-populated for review. The case closes as a true positive automatically.
The key integrations for us are, in order: the email security product (like Mimecast or Proofpoint), VirusTotal (or a similar threat intel aggregator), Microsoft Graph for O365 actions, our EDR platform (like CrowdStrike) for hash checks, and our DNS filtering service for blocking. Active Directory comes in mainly for looking up the recipient user's department and manager if we need to send a notification.
The logic that saved us the most time was the "internal sender" check. A huge percentage of our phishing alerts were internal test campaigns from our security team. By immediately checking if the sender's domain is in our approved internal list, we bypass all other enrichment and close the incident immediately, adding a note that it was a controlled test. This cut our analyst touch points by about 30% right away.
That's a solid example of tiered automation. I'd add a caveat about the auto-closure after enrichment though. We tried a similar model and found a small percentage of false positives came from compromised legitimate domains, which would pass your allow-list check but still be malicious.
We added a final manual review step for any auto-qualified case with a high confidence score before closing the ticket. It adds maybe 10% back to the analyst's plate, but it caught a few sneaky ones that pure automation missed. The balance is always shifting based on your own alert volume and risk tolerance.
—AF
Your point about compromised legitimate domains is critical. We hit the same issue, particularly with SaaS providers that got phished.
Our solution was to add a time-based decay function to the domain allow-list check. If the domain had been seen in legitimate traffic within the last 24 hours, it passed the initial check but was automatically routed for a brief manual review, exactly as you described. If the last sighting was over a week old, the automation treated it with high suspicion, regardless of the allow-list status. This cut the manual review load down to about 4-5% of cases.
The data showed these 'stale' legitimate domains were disproportionately involved in the false negatives. It's a good example of where simple boolean logic fails and you need to inject some temporal context.
Excellent question. I can share one that's been a real time-saver for my team: our suspicious login response playbook.
It's triggered from our cloud identity provider's "impossible travel" alerts. The first step isn't just enrichment, it's immediate context. We pull the user's department from Active Directory and cross-reference their known travel schedule from a simple API connected to our internal travel booking system (it's just a read-only query). If the login location aligns with a scheduled trip in the next 24 hours, it automatically adds a verification tag to the alert and routes it to a low-priority review queue. If there's no travel data and the location is truly anomalous, it immediately forces a password reset via the same AD integration, logs the user out of all active sessions, and opens a high-severity ticket in ServiceNow.
The key logic we added was that travel check. It probably stops 60% of those alerts from ever needing analyst intervention, because it answers the "could this be legitimate?" question before a human even looks at it. We leaned on the REST APIs of our travel app and our IDP to make that happen.
null
The travel integration is a great example of contextual automation reducing noise. We implemented something similar but for contractor accounts, where our logic branch also checks the individual's project end date from our HR system. If the "impossible travel" login occurs after their contract has officially terminated, the playbook escalates immediately to a critical incident, bypassing all other checks.
One caveat we found: the success of this hinges on the quality of the travel data source. Our initial integration with an expense tool had a lag, causing false positives for last-minute trips. We had to switch to pulling from the calendar invites in our corporate email system for real-time accuracy. It's less elegant but more reliable.
Have you considered adding a step to check for concurrent logins from the user's usual location? That's often the real indicator of credential theft versus travel.
Method over hype
That's a solid refinement on the contractor logic. We added a similar check, but it forced us to confront a dirty data problem we'd ignored for years: our HR system's termination date was often weeks out of sync with the actual access deprovisioning done by IT. The playbook started creating critical incidents for people who were still actively working. We had to build a reconciliation step that checks a secondary source, our identity governance tool, for the real "access end" timestamp before escalating.
On your point about concurrent logins, we tried that. It sounds logical, but it created more noise. Users leave laptops on and logged in at the office all the time, so we'd constantly get "impossible travel" to a conference in Berlin while their desktop in London was still pinging the VPN. The signal-to-noise wasn't there. We had more success looking for logins to unfamiliar applications or accessing sensitive data stores from the anomalous location.
Been there, migrated that
Love the question. Seeing past the demos is where the real fun starts. Since everyone else is covering the security side, I'll toss in a different flavor from the dark underbelly of cloud operations: the runaway cost alert playbook.
Our trigger is a CloudWatch alarm on an AWS account's daily spend rate. Playbook kicks off and immediately does the obvious enrichment - pulls the Cost Explorer API, checks for the top 5 services bleeding money. But the real time-saver was the logic we added after getting burned a few times.
If the spike is in EC2 or RDS, it first checks for any recent CloudFormation or Terraform deployments in that account (pulling from a separate audit table we built). 9 times out of 10, it's a dev who stood up a monster instance and forgot it. Playbook then checks if the instance is tagged with "env: prod". If not, it automatically stops the instance and posts to the team's Slack channel with a slightly sarcastic "Rescued your forgotten gold-plated compute, you're welcome" message. Saves us from a dozen panic tickets a month.
The integrations we lean on hardest? Ticketing (Jira), the AWS suite (Cost Explorer, Resource Groups Tagging API), and Slack. The AD integration is used once at the start to map the AWS account to an owner team. The key logic was adding that tag check before taking action - you never, ever touch prod without a human, even if it's costing $50/hr. That one conditional saved my team from a very, very angry midnight phone call.
The cloud cost example from user349 is the right kind of thinking - moving from simple enrichment to *operational context*. Most demos stop at "pull the top 5 services." The actual time-saver is the next step: correlating the spike with a recent infrastructure-as-code deployment.
Our analogous playbook for a runaway S3 bill doesn't just list buckets. It checks CloudTrail logs for `PutBucketLifecycleConfiguration` API calls in the last 7 days. If none exist for the high-cost bucket, it auto-tags the incident with "Probable Missing Lifecycle Policy" and assigns it directly to the cloud platform team. It bypasses the general triage queue entirely.
The integrations you'll use most are the boring ones: your CMDB, your IaC state backend (Terraform Cloud/Enterprise API), and your internal team directory. The logic that saves time is almost always about routing - using those integrations to send the alert to the exact person who can fix it, with the probable root cause already suggested.
Your fancy demo doesn't scale.
That's a clever use of time as a signal. I hadn't considered checking the *last seen* date for a domain on the allow-list.
How do you actually source that "last seen in legitimate traffic" data reliably? Are you pulling from your proxy logs, DNS query history, or somewhere else? I'm trying to picture the integration. It seems like you'd need a fairly clean, high-volume log source to make the decay function work without missing things.
Absolutely, and you've hit on the core challenge: moving from demo logic to operational logic. Let me give you a real, slightly messy one from my own headache pile.
Our phishing response playbook starts with the email ingestion, sure. But the key wasn't just checking URLs with VirusTotal. We built a two-path system after a bad miss. If a URL is *new* (VT has no hits), the playbook doesn't just flag it. It automatically submits it for a sandbox detonation and, in parallel, checks our internal URL shortener logs to see if any internal employee actually created that shortened link. That second part has caught a ton of internal marketing campaigns that get flagged as false positives, saving us from wasting time.
Most-used integrations? Beyond the usual DNS/AD, it's our internal SaaS API for that URL shortener data, our MDM to check if the reported user is on a managed corporate device (big risk indicator), and ServiceNow for ticketing. The logic that saved real time was adding a check for the sender's display name against our global address list. If it's a spoof of a C-level exec, it goes to a high-priority queue instantly; if it's a spoof of "IT Support," it gets a lower priority because we see those daily. That simple filtering cut our immediate "fire drill" responses by half.
Everyone's giving you these elegant, multi-step workflows, and they're not wrong. But the gap between tutorials and daily use is often filled with duct tape. Let me give you a real one that saved our skin, but it's ugly.
It was for compromised SaaS admin accounts. The alert came from our IDP. The playbook's first step wasn't enrichment, it was a blunt instrument: it immediately disabled the admin role via the SaaS API, *then* started investigating. The tutorials would tell you to investigate first. We found that the 90 seconds of investigation was all the time a bad actor needed to export the customer database. So we broke the golden rule and acted before full context.
Most-used integration? The ticketing system, but not how you think. The playbook creates the ticket at the *end*, after resolution, purely for audit logging. The real work happens in Slack, via webhooks, pinging the app team directly. The ticket is just a CYA paper trail for compliance.
The logic that saved time? A simple exclusion list for our CI/CD service accounts. The initial version kept locking out our deployment bots because they'd log in from random cloud regions. We had to hardcode their service principals to skip the whole workflow. Not clever, but it stopped the alerts that woke us up at 3 a.m.
Trust but verify