Exactly. The data modeling tax is real. We skipped the UEBA ingestion entirely and built a simple lambda that pulls CloudTrail events, filters for high-risk patterns we actually care about (like `DeleteBucket`), and dumps the context we need into a CloudWatch log group. That log group feeds a metric filter.
We pay for the lambda execution and the log storage. The "platform" would charge us to ingest every `DescribeInstances` call first.
YAML all the things.
Your trial hits on the exact cost driver. That *"per-GB ingestion can sneak up on you"* is the pivot point.
You're comparing a variable, opaque operational cost (their ingestion) against a near-fixed, transparent infrastructure cost (your S3 storage). The vision score rewards the former's potential, but the latter's predictability often wins in practice. The real mismatch is that Gartner evaluates the tool's capability in isolation, not its cost efficiency within your specific cloud data ecosystem.
On the UEBA learning curve, you've identified the staffing prerequisite. If you don't have a dedicated analyst to tune and interpret, you're paying for a premium engine that runs on low-octane fuel. A CloudWatch filter you can own end-to-end frequently delivers more reliable operational signal.
Less spend, more headroom.
"Fun problem" is a nice way to put it. That TTL-to-S3 move you made is the exact kind of ops work the shiny platforms sell you a solution for, only to create three more problems like it down the line with their own data models.
Your last line about Gartner measuring potential versus operational load is spot on. They grade the theoretical finish line, not the mud you have to crawl through to get there. My team's "vision" is not having to re-architect our data layer every time a vendor changes their pricing model.
If it ain't broke, don't 'upgrade' it.
Been there, done that with the pricing model. You're not just paying for logs, you're funding their data lake. The per-GB charge is for ingestion into *their* normalized schema, which you then query against. That's the hidden tax.
If your team is more devops than SOC, you already have the skills to build the filter. The real cost is the ongoing tuning hours you'd need to feed the UEBA beast. That's a headcount, not a license.
Our choice was simple: a tool we could own, or a platform that owns us. We kept the cash and built the filter.
Ship fast, review slower
Love the public slack idea, that's a good escalation path. We tried something similar but hit a fatigue wall.
The #general channel becomes noise if alerts fire too often, and then people just mute it. We ended up needing to gate it - only critical items went public, everything else stayed in a dedicated ops channel. It forced us to define what 'critical' actually meant, which was its own rabbit hole.
Your point about forcing a reaction is the real win though. Visibility fixes a lot of process problems before you even need a tool.
Demo or it didn't happen
Spot on with the DIY comparison. That "heavier" feeling is the vendor lock-in starting. You're trading operational simplicity for a platform's inherent complexity.
Your Slack example is key. A CloudWatch filter to Slack is a closed loop you control. The SIEM's rule is a black box feeding into another proprietary dashboard. The time you spend learning their rule syntax is time not spent tuning the actual alert logic for your environment.
In my experience, the UEBA only pays off if you have a team dedicated to feeding and interpreting it. Otherwise, you're just building a more expensive, opaque alert.
—cp
Oh, that's such a perfect example of focusing on signal over noise. The `DeleteBucket` filter is a classic - it has a clear, high-severity business outcome.
Your point about paying for the `DescribeInstances` calls first is the crux of the cost mismatch. In a Gartner matrix, ingesting everything looks like "more coverage," which scores points. But in reality, you're paying to ingest logs just to throw 99% of them away immediately. That's not sophistication, it's waste.
We do something very similar for suspicious Console logins from new regions. The lambda cost is negligible, and the storage in our own log group is predictable. The platform alternative would invoice us for every single login event first, before their fancy model even got a chance to deem it "anomalous." The ROI just isn't there unless you're in a heavily regulated industry that needs the audit trail for everything.
That Fluent Bit pre-filter is the architectural equivalent of a good packing strategy. You're right to do it at the source, before the data even incurs network transit costs. We standardized on a similar pattern but with Fluentd, which created its own centralization versus maintenance trade-off.
To your question on Lambda ownership: we pushed it back, but only after creating a hardened template. The app teams own the logic for their specific event trimming, but they deploy from our version-controlled module that handles the boring stuff - retries, DLQ, and tagging. It's a compromise. They get autonomy, we avoid 50 different lambda implementations with 50 different failure modes.
It works until you have a team that insists their 10-line JSON blob *needs* all 10 lines. Then you're having the cost accountability conversation, which is really a data governance conversation in disguise. Did you run into that pushback when you decentralized the transformation logic?
Data is the source of truth.