Hey folks! 👋 I've been diving deep into Cortex XSIAM for the past six months, moving our startup's security monitoring from a more traditional SIEM+SOAR stack. The promise of the "AI-driven, single platform" for everything from ingestion to automated response is super compelling, especially for a lean team.
We're now at the point where we're considering a true 24/7 SOC model, but I'm really curious about real-world experiences. We're currently covering about 18/5 with on-call for nights/weekends, and the jump to 24/7 feels like a big stepβnot just in staffing but in how we leverage the tool itself.
* **Staffing:** For those running 24/7, what does your analyst tier structure look like? How much are you relying on L1 analysts working primarily within XSIAM's incident interface vs. needing deeper investigation skills? Has the AI/automation actually reduced the headcount needed per shift, or just changed the skill set?
* **Tooling & Playbooks:** XSIAM's built-in automations and playbooks are great for common stuff. But for 24/7 coverage, we've found we need even more robust "shift handoff" and "context carryover" features. Have you heavily customized the incident lifecycle or built custom widgets/playbooks for shift operations?
* **The Cost/Value Trade-off:** We're watching our data ingestion like a hawk. Moving to 24/7 monitoring means more proactive hunting and potentially more logs. Any hard-won lessons on keeping XSIAM costs predictable while ensuring analysts have what they need?
I'd love to hear your lessons learnedβboth the wins and the "watch-outs." Our goal is to build an efficient, alert-fatigue-reduced SOC that actually leverages the platform's strengths, not just fights it.
That jump from 18/5 with on-call to 24/7 is exactly where we were a year ago. On your staffing question, the skill set definitely shifted more than the headcount. We found we could run with lighter L1 staffing on the overnight shifts specifically because XSIAM's correlation and automated enrichment is pretty good. But - and it's a big but - those analysts need to be really sharp at reading the automation's work and knowing when to escalate or break out of a playbook. A sleepy L1 just clicking "run playbook" won't cut it.
For your point on shift handoff, absolutely critical. We built a custom dashboard widget that pulls in all "in-progress" incidents and pins the last shift's notes/comments right at the top. The out-of-the-box lifecycle didn't quite capture the nuances of a true handover, like a hunch an analyst was following. We also ended up writing a few custom playbooks just for shift-change procedures, like auto-assigning and sending a summary digest to the incoming lead.
Has the AI actually reduced headcount? For us, it let us reallocate bodies from manual triage to more threat hunting and playbook tuning. But you still need warm bodies in the chair for those 3 AM critical alerts.
Data nerd out
Moving from on-call to 24/7 coverage really changes how you interact with the automation. We saw something similar - the playbooks handle maybe 70% of common alerts cleanly overnight. But that remaining 30% needs someone who can recognize when the automation is stuck in a loop or has made a weird enrichment call. It's less about headcount and more about having at least one person per shift who can think outside the XSIAM interface box.
For shift handoff, we also had to get creative. The out-of-the-box incident timeline wasn't cutting it. We ended up using the API to push a summary of all open cases into a Slack channel at the top of each shift, with a direct link back. It's a duct-tape solution, but it works better than expecting everyone to remember to update a custom field.
That context carryover is a real pain point, especially for slower-burn incidents that span multiple shifts. Did you guys build any custom dashboards or widgets for that, or are you mostly relying on manual notes?
ship it
> The out-of-the-box incident timeline wasn't cutting it.
Yeah, we ran into the exact same thing when we started testing overnight shifts. That manual Slack idea is clever though. We haven't built any custom dashboards yet, mostly because I'm still getting my head around the Cortex API.
I'm curious, when you say you need someone who can "think outside the XSIAM interface box" for that 30%, what does that actually look like on a night shift? Are they jumping into the AWS console or raw logs directly, or is it more about overriding a playbook decision? Trying to picture the actual workflow.
Still learning
Great question. We're in a similar boat, trying to plan for the same transition. The staffing part is tricky because, like others said, the skill set changes.
On the tooling side, we've found the built-in incident lifecycle a bit rigid for handoffs too. We're looking at using the comment system almost like a log, but it gets messy. Curious if anyone's found a clean way to build a "shift summary" view right into the main incident screen without a ton of API work?