Hey everyone, been lurking here for a while but finally have a reason to post! 😊
Our team is in the middle of a major security data consolidation project. We're a ~200-person company running almost entirely on AWS (EC2, ECS, some Lambda). We've built a bunch of custom Python-based data pipelines (mostly using **boto3**, **requests**, and some **Apache Airflow** for orchestration) that pull logs from various SaaS apps, our own product, and cloud services. Currently, they just dump JSON parquet files into an S3-based data lake. It works for analytics, but our security folks are rightfully asking for a proper SIEM (and eventually SOAR) for correlation, alerting, and investigation.
We're evaluating options and I'm hitting that classic "build vs. buy vs. use a managed service" dilemma. Our key constraints and desires are:
* **Ingestion Volume:** Currently about 45 GB/day of log data, but expecting to double as we onboard more sources.
* **Existing Investment:** We don't want to throw away our Python pipelines. We need a SIEM that plays nice with custom ingestion, ideally via API or a flexible agent/collector.
* **AWS-Native Preference:** We'd love to stay within the AWS ecosystem if it makes sense, but aren't opposed to a best-of-breed third party.
* **Future SOAR:** We want a path to automation and playbooks, even if we start with just the SIEM.
So my core question is: **For a shop like ours, with existing Python/AWS pipeline expertise, what's the most efficient SIEM to adopt?**
I've been prototyping a few approaches, and here's a snippet of the kind of pipeline we'd be looking to integrate. This one fetches Okta logs and pushes them... somewhere.
```python
import boto3
import requests
import pandas as pd
from datetime import datetime, timedelta
def fetch_okta_events(api_key, last_hours=1):
"""Fetch Okta system logs for the last N hours."""
headers = {'Authorization': f'SSWS {api_key}'}
url = f'https://{tenant}.okta.com/api/v1/logs'
params = {
'since': (datetime.utcnow() - timedelta(hours=last_hours)).isoformat() + 'Z',
'limit': 1000
}
response = requests.get(url, headers=headers, params=params)
response.raise_for_status()
return response.json()
# This is where the SIEM integration would happen!
# Currently we write to S3, but we'd need to transform and send to the SIEM's ingestion endpoint.
def send_to_siem(event_batch, siem_ingestion_url, siem_key):
# Would this be a HTTP/S POST with a specific schema? A syslog forward? An agent SDK?
pass
```
The options we're currently pondering:
* **AWS Native:** Amazon Security Lake with a third-party SIEM on top? Or straight to Amazon OpenSearch Service with the Security Analytics plugin? This feels clean but I'm worried about the detection engineering and management overhead.
* **Cloud-Native SIEMs:** Looking at providers like Panther or Sumo Logic. They seem to embrace programmatic, pipeline-friendly ingestion and are built for cloud scale.
* **Traditional Heavyweights:** Splunk, Sentinel. The power is undeniable, but I'm concerned about cost spirals and whether they'll fight our custom pipelines instead of complementing them.
What I'd love to hear from the community:
* Any horror stories or success tales from integrating custom Python pipelines with a specific SIEM's ingestion API?
* For a team of our size, is managing something like OpenSearch Security Analytics a full-time job?
* Are any of the newer cloud SIEMs truly scalable and cost-predictable when you control the ingestion pipelines?
Really excited to learn from your experiences. The data flow and integration piece is critical for us, so the "plumbing" might be the deciding factor.
Data nerd out
Data nerd out
I'm a PM at a 150-person tech consultancy where we manage client cloud deployments, and we went through this evaluation last year. We run a similar AWS/Python stack and ultimately deployed **Elastic SIEM** on EC2.
**Core comparison for your scenario:**
**Cost Control**: Elastic's self-managed model was crucial for us. We pay for EC2 infra (~$2k/month for a 3-node hot-warm setup) plus the licensing, and ingest cost doesn't spike unpredictably. Managed cloud SIEMs (Splunk Cloud, Sentinel) quoted us based on daily GB, which would have put us at $6-9k/month at your volume with expected growth.
**Custom Pipeline Integration**: Elastic's HTTP Event Collector (HEC) is a simple JSON-over-HTTPS endpoint. We pointed our existing Python pipelines at it with minor tweaks to the final POST step; no agent replacement needed. Sentinel's Data Collector API felt more opinionated and required more adjustment.
**AWS-Native Fit**: Sentinel is the obvious deep AWS integrator for alerting on CloudTrail, GuardDuty, etc., and we use it for some Microsoft-centric clients. But for a custom data lake, Elastic's beats agents and the S3 input plugin let us also read existing S3 parquet files directly, which avoided a dual-pipeline rewrite.
**Operational Overhead**: The honest limitation. Running your own Elastic cluster means you own indexing strategy, node failures, and version upgrades. We had a devops engineer spend 2-3 days/month on maintenance. A true cloud SIEM eliminates that entirely.
**My pick:** I'd recommend the self-managed Elastic SIEM stack for your case, specifically because you have existing Python pipelines and want to keep direct control over ingestion cost and data flow. If your team has zero appetite for managing the underlying search cluster, tell us your devops headcount, and we should compare managed services instead.
You're right to focus on the integration with your Python pipelines. That's the linchpin.
While many SIEMs have generic HTTP collectors, the real test is how they handle the semi-structured, nested JSON that usually comes out of these custom pipelines. Some systems require rigid, flattened CEF formatting, which would force a major rewrite of your final transformation stage. Others, like the Elastic stack user805 mentioned, are more forgiving with native JSON, letting you map fields later.
Given your AWS-native preference, have you looked at **AWS OpenSearch Service** with the Security Analytics plugin? It's the managed version of what they're describing. The built-in S3 ingestion (via S3 Scan) could be a path for your existing parquet files, while your live pipelines could POST to an HTTP collection endpoint. It keeps you within the AWS console and IAM policy ecosystem, which might simplify your team's operational overhead compared to managing EC2 instances.
benchmark or bust
OpenSearch Service's S3 ingestion is painfully slow for anything beyond initial backfills. It's a batch process, not a real-time pipeline.
You'd still need that HTTP endpoint for live data from your Python scripts, which makes the S3 path redundant. I'd stick with a direct HEC/HTTP ingest from Airflow and treat the data lake as a cold archive.
Also, the Security Analytics plugin is far behind Elastic's detection rules. You'll spend more time writing custom rules than you save on instance management.
—cp
You mentioned a 200-user AWS shop with Python pipelines. That's close to our size! We run a similar setup.
How are you handling IAM roles for your pipelines to talk to a SIEM endpoint? I'm trying to avoid embedding keys in our Airflow tasks. Would you use instance profiles or something like a Lambda with a VPC endpoint? Still figuring out that security piece myself 😅
Your build vs. buy vs. managed dilemma is the core challenge. With your existing Python pipelines and AWS-native preference, the managed services become a lot less attractive due to vendor lock-in and inflexible pricing for custom data shapes.
You'll get the most mileage from a solution that treats your pipelines as a first-class citizen, not an afterthought. I'd focus your evaluation on the schema flexibility of the ingestion endpoint. A system requiring strict CEF or a proprietary agent will force a total pipeline rewrite, while one accepting raw JSON with dynamic mapping will let you reuse 90% of your code.
Given your volume and expected growth, the pricing model is just as critical as the tech. Have you calculated the 3-year total cost of ownership for a managed SIEM at 90+ GB/day versus running the compute yourself on reserved EC2 instances? The delta can fund an entire security engineer.
Data > opinions