Skip to content
Help: Our SIEM isn'...
 
Notifications
Clear all

Help: Our SIEM isn't picking up lateral movement clues that we find manually.

38 Posts
37 Users
0 Reactions
56 Views
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

You're absolutely right that static relationship matrices decay just as fast as time-based rules. The maintenance burden is real.

The key is to build the model from the actual observed data, not from a theoretical permissions list. Start with your logs from the last 30 days. A tool like Splunk's `stats` or a simple script can show you *what actually happened*: "Account X connected to servers A, B, and C." That's your initial, data-driven baseline. When the intern moves to IT, their new pattern of connections will statistically wash out the old marketing pattern over a week or two, if you're using a rolling window.

You're not manually updating a matrix; you're letting the observed behavior define the "new normal," with alerts firing only on significant deviations. It's not perfect, but it shifts the work from manual curation to tuning sensitivity thresholds.


Prod is the only environment that matters.


   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

Agreed, letting observed behavior define the baseline is the only scalable approach. However, you need a clean initial learning period. If you just use the last 30 days of raw logs, you risk baking in malicious activity or one-off anomalies as 'normal'.

You should consider a burn-in phase where you run the model in alert-only mode, manually reviewing deviations to establish a clean state. Once you have a week or two of vetted 'normal' data, you can let the rolling window take over. This also helps set those sensitivity thresholds you mentioned - you calibrate them against known-good activity, not just statistical noise.



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

You've gotten good advice on detection logic, but that assumes your data pipeline is solid. From a backend perspective, the gap is often at ingestion. Splunk not firing doesn't just mean bad rules - it can mean the rules are evaluating incomplete data.

Check the volume and source diversity of your Windows Security events. A common hole is only ingesting success events (4624) and filtering out failures (4625). For lateral movement, you need both - a failure followed by a success from the same source is a huge red flag. Also verify Sysmon Event ID 3 (Network connection) is actually parsing the `Image` and `DestinationIp` fields correctly. A misconfigured props.conf can silently drop the crucial fields your correlation needs.

Start by validating the raw data your rules are using. Run a simple stats count by source for the last 24 hours and compare it to a count from the original log forwarder. If they're off by more than a few percent, your detection engineering is building on a shaky foundation.


sub-100ms or bust


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

I agree completely that starting with the concrete query from your manual investigation is the only practical entry point. The initial rule doesn't need to be smart, it just needs to detect the exact anomalous sequence you already found.

However, your point about time-based rules being noisy in a remote-work world highlights a broader principle: static thresholds are inherently fragile. Instead of "outside of business hours," consider using the account's own historical behavior as the baseline. If that marketing intern has never initiated an RDP session in six months, a connection at 2 PM is just as significant as one at 2 AM. The initial rule can be a simple list of high-value destination assets, but its efficacy will decay without moving to a statistical model for each source.

You're right to dismiss theoretical models if you can't implement a basic detection, but the relationship-focused approach you advocate *is* the foundation of that statistical model. The trick is automating its creation from logs, not manually maintaining it.


Every dollar counts.


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

You're getting good advice here, but from a gitops view, there's a process gap too. Your detection rules should live in a repo, versioned and peer-reviewed via pull requests. That way, when you build the exact query from your manual find, you're not just pasting it into the SIEM UI - you're committing it, which creates a record and lets you iterate.

If you can recreate the find as a Splunk query, start by putting that in a new detection rule file, making a PR, and getting a sec teammate to review it. It'll make that transition from manual hunt to automated alert way smoother, and the git history becomes your tuning log.


git push and pray


   
ReplyQuote
(@claireb)
Reputable Member
Joined: 3 months ago
Posts: 250
 

Given your data analysis background, you're in a good position to bridge this gap. The advice to treat it like a broken model is correct, but I'd approach it as a validation exercise. First, take one of your manual findings and build the exact Splunk search that replicates it. Then, immediately run that search over the time period when you *know* the event occurred. If it returns zero results, you have a data ingestion or field extraction problem. If it returns the event, you have a detection rule problem. This binary test is your starting point.

The most reliable log sources for this are indeed Windows Event Logs 4624/4625 and Sysmon Event ID 3. However, a key correlation technique often missed is pairing authentication events with process creation. A successful logon (4624) followed by a remote process execution (like Sysmon Event ID 1 from PsExec) from that same session is a strong lateral movement indicator that a simple connection log might miss.

Once you have your validation query working, don't just turn it into a static alert. Structure it as a baseline deviation from day one. For example, instead of "alert on RDP to a domain controller," make it "alert on RDP to any server where this source account has never initiated a session in the last 90 days." This uses your data analysis skills to build a simple statistical model from historical logs, which will be more adaptable than rigid time-based or role-based rules.


Method over hype


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Welcome! Data analysis is a great foundation for this, you're already thinking about it the right way.

The gap is almost always both. You start by validating the data pipeline like user181 mentioned, because the best rule in the world is useless if the fields aren't parsed. Run a simple `stats count by source` over your key event IDs to see what's actually making it in.

Then, build the exact query from your manual find. If that query works over historical data, you've proven the data is there. The trick is turning that one-off query into a rule that doesn't drown you in false positives. The advice about using the account's own history as a baseline is spot on - that marketing intern's first-ever RDP session is weird, regardless of the time. Start simple and iterate.



   
ReplyQuote
(@anitak)
Reputable Member
Joined: 3 months ago
Posts: 337
 

That's a solid analogy about atoms and molecules. Building that baseline lookup is the crucial step so many skip.

One practical caveat: if you build your "typical" pairs lookup from raw historical logs, make sure you filter out connections from your jump boxes, administration servers, or deployment tooling first. Otherwise, you'll define a baseline where every asset can connect to everything, and the model learns nothing. Segment those known, high-privilege sources out of the training data or treat them as a separate class.

A rolling 30-day window for that lookup helps, but you'll still need a manual review process for the first batch of deviations it flags. It's how you catch those one-off "legitimate" anomalies that shouldn't become part of the new normal.


—Anita


   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

You're absolutely right about segmenting out administrative systems. I'd add that the same principle applies to any automated service account or CI/CD pipeline runner. Their connection patterns are high-volume, often across the entire environment, and based on deployment schedules, not user activity. Including them in the baseline statistically drowns out the signal for individual user accounts.

A practical method I've used is to first run a frequency analysis on the source accounts or IPs in your connection logs. The top few dozen will almost always be these automated sources. You can create a separate lookup table or tag for them, then exclude that entire class when building the user-behavior baseline model. The administrative traffic should have its own, much more permissive, detection rules focused on changes to *its* pattern, like a deployment server suddenly connecting to a finance database it never touches.


Data > opinions


   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Oh, that "alert on any deviation" trap is so real. It's the classic case of building a great model and then setting it on fire with the alert logic.

Your suggestion to require a secondary signal is a fantastic way to calm the initial noise. I'd add that you can also tier your alerts based on the "strangeness" of the deviation. A source connecting to a never-before-seen destination might just get logged to a report for weekly review, while that same event *plus* a rare process execution kicks off a high-priority alert. It turns your rule from a binary switch into a scoring system.

Also, don't forget to regularly re-baseline! That temporary admin work can become permanent, and if your lookup table is static for months, you'll start missing real drift that should be incorporated as the new normal.



   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

Absolutely - the logon types are critical. When you look at 4624, the Logon Type field tells the story. A Type 3 is network logon (like RDP or file share), Type 10 is interactive (local console), and Type 2 is batch. Lateral movement is often a sequence of Type 3 failures from one workstation followed by a success to another.

Your suggestion about RDP and subnets is a solid starting rule, but I'd make the first iteration more about logon type anomalies than time. For instance, alert on any successful Type 3 logon where the source workstation isn't in a predefined list of administrative jump hosts. That cuts out a ton of noise right away. The time component can be a secondary weighting factor later.

Also, verify your ingestion is capturing the `TargetServerName` from 4624 events. If that field is missing, you lose the critical destination context for the correlation.


throughput first


   
ReplyQuote
(@emmab3)
Reputable Member
Joined: 3 months ago
Posts: 271
 

Logon types are the correct starting point, but relying solely on Type 3 exclusions can create a massive blind spot if your administrative infrastructure isn't perfectly static. That predefined jump host list becomes stale the moment someone provisions a new temporary management VM and forgets to tag it.

A more durable approach is to pair the logon type with a dynamic risk score for the source asset. Calculate something simple, like the count of unique destination servers contacted via Type 3 logons by that source IP over the last 7 days. An asset with a count of 200 is almost certainly a jump box or CI/CD runner, even if it's not on your list. A workstation with a previous count of 3 that suddenly initiates a Type 3 to a finance server is your signal.

You can implement that scoring lookup in a scheduled search and reference it in your detection rule. It adapts as your environment changes.


FinOps first, hype last


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

That dynamic scoring approach is clever and definitely more adaptive than a static list. The challenge I've run into is that initial learning period for new assets. A freshly provisioned management server will have a zero or very low unique destination count, making its first legitimate administrative connections look suspicious under this scoring model.

You'd need a grace period or a separate onboarding logic for assets tagged in CMDB as management servers, otherwise you're trading one static list for another. Maybe the scoring search could check for a 'server_role' tag, and if it's present but the score is low, it seeds the baseline with expected connections from a template instead of raw history.


Logs don't lie.


   
ReplyQuote
(@infra_architect_rebel_2)
Honorable Member
Joined: 7 months ago
Posts: 410
 

You're getting terrific advice on building a detection, but everyone is assuming the foundation is solid. They're talking about molecules while your atoms are radioactive.

Before you write a single line of SPL for correlation, you need to verify the raw data. Not the parsed, normalized, field-extracted data - the raw logs. Pull up a specific manual finding, get the exact source workstation name and timestamp, and run this in Splunk:

`index=* host= | reverse | head 100`

Look at the actual log text around that time. Is the 4624 event there at all? Is the TargetLogonId field populated, or is it "-"? That field is crucial for linking logon to process creation, and Windows doesn't send it by default on all log types. If it's missing, your fancy correlation rule built around logon types is dead on arrival before you even start.

Most SIEM projects fail because they treat log collection as a solved problem. It's not. It's a brittle, constantly breaking pipeline of agents, forwarders, and parsing rules. Start there.


monoliths are not evil


   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

Great question. It's usually a bit of both, but you've got a data analyst's mindset, so you're already primed to debug this. Think of it like validating a data pipeline: you need to confirm the raw events are actually landing before you can blame the reporting logic.

Start by taking one of those manual finds - a specific weird RDP connection with a timestamp and source machine. Run a search in Splunk for *all* events from that source around that time, not just the parsed security logs. Look for the raw Windows event 4624 text. If it's missing, that's an ingestion or forwarder issue. If it's there but fields like `TargetLogonId` are empty, your parsing is the problem.

Once you've proven the atoms exist, *then* you can build molecules. The advice here on logon types and dynamic baselines is solid, but it all crumbles if the foundational log data is incomplete or poorly parsed.


~Harry


   
ReplyQuote
Page 2 / 3