We're building out a new SOC and need to finalize our SOAR platform. Budget is a factor, but so is team size and maintenance overhead.
Has anyone run Shuffle in production at scale? I'm concerned about:
* Custom integration development effort
* Long-term playbook maintenance cost
* Scaling the workflow engine under alert load
Cortex's vendor-supported integrations and XSOAR features look good, but the price tag is steep. For those who chose it:
* What specific features justified the cost?
* Did it actually reduce your mean time to respond (MTTR) compared to your POC with open source?
I manage everything with Terraform, so I'm also looking at the infrastructure footprint and deployment complexity for both options.
—cp
I'm a security engineer at a mid-market SaaS company, around 500 employees, and we've been running Cortex XSOAR in production for our SOC for about 18 months after evaluating Shuffle and a couple of other options.
* **Team size and maintenance burden:** This was the decider for us. With a SOC team of 8 people, we couldn't spare the 0.5-1 FTE of engineering time needed to keep Shuffle's integrations updated and maintain the core platform. Cortex's library of 1,000+ vendor-supported integrations means our analysts write playbooks, not API clients. We spend about 4 hours a week on playbook maintenance vs. the estimated 2 days we budgeted for Shuffle upkeep.
* **Real pricing and TCO:** Cortex's sticker shock is real, starting around $35k-$40k annually for a basic pack, but that's for the platform and unlimited integrations. The hidden cost of Shuffle is your team's salary time. For us, the engineering hours needed to build and maintain custom connectors for our critical tools (like our EDR and internal ticketing) would have burned through that price difference in less than two quarters.
* **Deployment and infrastructure footprint:** Both deploy pretty cleanly with Terraform. Cortex runs as a set of VMs/containers and we manage it on our own Kubernetes cluster. It's not lightweight, requiring about 8 vCPUs and 32GB RAM per node for comfortable performance. Shuffle's footprint is smaller, but scaling the workflow engine under load required us to tune the Celery workers and Redis configuration in our POC, adding complexity we didn't want to own.
* **Where the open-source option hits a wall:** In our 90-day POC, Shuffle handled our baseline alert load of ~500/day fine. The breakdown came during a simulated incident surge. Pushing to ~2,000 alerts in a 4-hour window caused playbook execution delays and required manual intervention. Cortex's engine, in the same test, queued and processed the spike without degradation. The difference was in execution parallelism and managed message queuing.
My pick is Cortex XSOAR if you have a SOC team under 15 people without dedicated SOAR developer bandwidth. It's a premium for a finished product so your analysts can be analysts. If your team has a developer who can own Shuffle full-time for the first 6-12 months, and your alert volume is consistently under 1k/day, then Shuffle's path makes sense. For a clean call, tell us your actual team size dedicated to SOAR and your peak daily alert volume from the last month.
APIs > promises