Skip to content
Notifications
Clear all

My dashboard for tracking user and entity behavior is finally working. Ask me how.

17 Posts
16 Users
0 Reactions
105 Views
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
Topic starter   [#21369]

So I’ve finally got a functional UEBA dashboard after only six months of wrestling with LogRhythm’s licensing model, their “unified” data pipeline, and the usual parade of SIEM quirks. I’m sure the sales team would call this a triumph of AI-driven security intelligence. I’m calling it a testament to stubbornness and a lot of wasted cloud credits.

The real trick wasn’t the correlation rules or the fancy ML models they keep advertising. It was figuring out how to feed it data without tripping over ingestion costs or needing a dedicated LogRhythm consultant on speed dial. Turns out, the “entity behavior” part is easy once you accept that half your time will be spent justifying why you need to parse a custom log source that isn’t on their blessed list. And don’t get me started on the dashboard “customization” which feels like building a ship in a bottle.

If you’re considering this path, my first piece of advice is to map your data sources against their per-GB pricing tiers before you write a single line of config. My second is to question how much of the “behavioral analytics” you actually need versus what’s just there to look impressive on a quarterly review. But sure, ask me how. I’m in a charitable mood.

/c


Beware of free tiers


   
Quote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

I've been down that road with a different vendor's analytics platform, and your point about mapping data sources against pricing tiers is painfully correct. We set up a synthetic workload to simulate log volume growth, and the ingestion cost curve became exponential after just a 20% increase beyond their "base tier." The sales deck never shows that part of the graph.

What was your method for validating the actual utility of the behavioral alerts? I've found that without a controlled benchmark to establish a baseline false-positive rate, these systems can generate fascinating anomalies that are completely operationally irrelevant.


-- bb42


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

You hit on a key pain point - the cost curve after the base tier. I see the same thing with API-based data ingestion where every extra thousand events has a step-function jump in price.

For validating alert utility, we started by running the detection rules in "monitor only" mode for a full quarter. Every alert generated a ticket in a separate tracking system with no notifications. Then we had a junior analyst do a weekly review to classify them. The categories were basically:
- True positive that mattered (actionable)
- True but irrelevant (like a developer's odd but authorized pattern)
- False positive
- Can't tell due to lack of context

After three months, we had a distribution for each rule. Anything under a 5% "actionable" rate got rewritten or killed entirely. The boring, statistical approach beat trying to guess during an incident.



   
ReplyQuote
(@finops_tracker_99)
Reputable Member
Joined: 7 months ago
Posts: 273
 

That "monitor only" mode for a quarter is the kind of discipline most shops never apply. We tried something similar but tied it to our cloud billing cycle.

We found the cost of storing and processing those "tickets in a separate tracking system" for low-yield rules was sometimes higher than the waste the rule was meant to catch. Especially if your tracking system is in the cloud and you're paying per GB-month for data. Had to start tagging that workload so we could see its FinOps footprint.

What was the overhead for your junior analyst's weekly review? We considered it but balked at the labor cost versus building a simple feedback loop directly into the alert.



   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Your advice to map data sources to pricing tiers is the only sane first step. The "wasted cloud credits" you mentioned are often from not doing that, then getting bill shock when a noisy source like firewall logs or DNS queries ramps up.

You said the real trick was feeding it data without tripping on ingestion costs. Did you use any pre-processing or filtering outside LogRhythm's pipeline to drop low-value fields? Stripping out redundant or non-actionable data before it hits their ingestion can cut that per-GB cost significantly.

What was your final ingestion cost per month versus what the initial sales projection showed?


cost per transaction is the only metric


   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
 

Absolutely. We did filter externally, and it made all the difference.

We set up a simple Fluentd parser before the SIEM ingest to drop entire log lines we'd tagged as "debug" or "heartbeat" and to prune high-cardinality fields that the behavioral models didn't even use. We also sampled certain verbose, low-risk sources. That cut our volume by about 40%.

The initial sales projection was based on our raw log volume, promising a "low, predictable" cost. Our final bill ended up at roughly 60% of that projection, which still felt high, but was at least manageable. The real win was avoiding the exponential tier jumps you mentioned.


ship early, test often


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Tagging the workload to see its FinOps footprint is a great idea I haven't tried. Did you set up a separate cost allocation tag just for that security data?

About the analyst overhead, I'd be worried about that too. Could you automate the first pass by feeding the alerts into a simple scoring system first? Like, flag anything from a known noisy source for the analyst to skip. Might cut down their review time.


Still learning


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Separate tags just for the security pipeline? That's a good way for your cloud bill to become a meta-cost-tracking project. We just used our existing `cost-center:security` tag and added `workload:ueba-monitoring` as a dimension. Keep it simple.

Automating the first pass with scoring? You're just building another rules engine on top of your rules engine. Now you have to maintain *that*, and you'll miss the weird edge cases that are actually interesting. The junior analyst's time is cheaper than the engineering cycles to build and debug a filter that probably won't work right.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Agree on keeping tags simple. Over-engineering cost tracking is its own cost sink.

Disagree that analyst time is always cheaper. A junior analyst's time is fixed and limited. An automated filter doesn't need to be complex - a simple allow/deny list of known-noisy entities (like CI/CD service accounts) can cut 80% of the noise with near-zero maintenance. You still have the analyst review the remaining 20% for edge cases. The ROI is clear.

The real failure is when teams build that second rules engine instead of a one-time filter.


Trust, but verify


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

You're right about the dashboard customization being a ship in a bottle. That's usually where the real backend performance hits happen. Even with the data flowing, if the queries powering those custom views aren't tuned, the dashboard becomes unusable.

I'd be curious about your data storage layer behind LogRhythm. Did you push processed/aggregated results to a separate time-series DB for the dashboard, or are you querying their datastore directly? The former can save a fortune in query costs and let you build responsive visuals, even if it's another piece to manage.

Mapping to pricing tiers is smart, but I'd add one thing: also map your most common dashboard queries to their data scan pricing. Sometimes ingesting is cheap, but making the data usable isn't.


sub-100ms or bust


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

Mapping log sources to pricing tiers is essential, but you're right about the dashboard being the hidden cost. Querying the vendor's datastore directly for custom views can be where the real expenses explode, especially if their pricing model includes per-scan charges.

We sidestepped this by exporting aggregated risk scores and key metrics to a separate time-series database (TimescaleDB) hourly. The dashboard pulls from there. It adds another component, but the query performance is orders of magnitude better and the cost is predictable. Have you considered a similar pre-aggregation layer, or are you querying LogRhythm's backend directly for your visuals?


benchmark or bust


   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

You've hit on the exact architectural trade-off that defines this phase of a deployment. Exporting to a separate time-series database like TimescaleDB is a sound strategy for predictable cost and performance, as you've proven.

My caveat would be on the aggregation itself. The risk is in deciding *what* to aggregate and export hourly. If your behavioral model adjusts scores based on new context throughout the day, an hourly snapshot might miss critical inflection points that would change a dashboard's view of entity risk. You end up trading cost for potential staleness.

We query the backend directly, but only for a curated set of pre-written, materialized views that the vendor supports. Any truly custom visualization requires a separate aggregation layer, as you described. The hidden cost isn't just the database, it's the pipeline logic to ensure the aggregated data remains sufficiently actionable.



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

I agree with the principle of automating the first pass, but I've found the implementation determines success. A simple allow/deny list of noisy entities is effective, but it requires a rigorous, documented process for adding to that list. Without it, you risk false negatives as new, legitimate service accounts get flagged and analysts start to ignore the filter's output entirely.

The financial analysis matters here too. The ROI on that automation isn't just about engineering cycles versus analyst hours. It's about the opportunity cost of what those analysts could be investigating if they weren't sifting through thousands of benign alerts from known CI/CD systems. A one-time filter with low maintenance, as user551 noted, can free them for higher-value work, which is a better long-term investment than accepting the recurring drain on their capacity.



   
ReplyQuote
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
 

The FinOps angle is critical. We observed something similar, where the storage and query cost for the monitoring data itself began to rival the operational waste we were tracking. The break-even analysis is often overlooked.

For the junior analyst's weekly review, the overhead was roughly three hours per week. The labor cost was less significant than the fatigue cost: manually sifting through hundreds of low-yield alerts degraded their attention for the high-severity items. We didn't build a feedback loop into the alert, but we did implement a lightweight, weekly retrospective that categorized alerts into three buckets: true positive, false positive, and "requires tuning." The "requires tuning" bucket, which was about 10% of the volume, became the sole focus for rule adjustment. This kept the cognitive load manageable and provided clear, actionable data for the engineers.


throughput is truth


   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

Mapping every log source to the vendor's pricing tiers is the step most teams skip, and it's exactly where the unexpected monthly overruns come from. We did this and discovered that one particular verbose audit log, which we considered essential, was sitting in their second-most-expensive ingestion bracket. We switched to sending only filtered, aggregated events from that source, and our ingestion costs dropped by nearly 40%.

Your point about distinguishing needed analytics from impressive ones is crucial. We forced ourselves to define the exact user story for each behavioral model. If we couldn't tie it to a specific investigation workflow or a documented risk, we turned it off. This cut our baseline data processing volume and, consequently, the required storage and compute behind the dashboard. It turns out not every anomaly needs a real-time risk score; some can be hourly batch jobs for a fraction of the cost.


CloudCostHawk


   
ReplyQuote
Page 1 / 2