Coming from data analysis, you're spot on to think of this as a pipeline issue. The gap is almost always detection logic, but you have to verify the data first.
Take one of your manual finds and treat it like a broken data point. Find the exact Windows event ID 4624 in Splunk around that timestamp, but look at the raw event text, not the parsed fields. You're checking if fields like `TargetLogonId` or `TargetServerName` are actually populated. If they're missing or defaulted, your parsing is stripping out the atoms you need to build the correlation molecules everyone's discussing.
Assuming the data is good, I've found success focusing on logon type anomalies paired with a simple frequency baseline. Forget static jump box lists; calculate how many unique destinations a source hits via Type 3 logons per week. A workstation that normally talks to 3 servers suddenly hitting a 4th is a much clearer signal than trying to define "odd hours" globally. It turns lateral movement into a statistical deviation, which fits your background.
✌️
Welcome to the community. Coming from data analysis is a huge advantage here - you're already thinking about the problem the right way.
You've gotten excellent advice, but I'll add one more checkpoint from the procurement side. Sometimes the gap isn't ingestion or rule logic, but licensing. Certain advanced security logs, like detailed process auditing or specific Windows Event IDs, require upgraded data tier licenses or add-ons from your vendor. It's worth checking if the specific log sources mentioned here are even included in your current Splunk license package. I've seen teams build perfect rules only to realize the critical fields are being dropped at ingestion because they're not paying for that data category.
Assuming you have the right data, start with the simple frequency analysis that user621 mentioned. It's a great first pass that often reveals noise you didn't know was there.
Trust the data, not the demo.
Totally agree on separating service and user accounts. That's a huge timesaver.
For user baselines, the 30-day window is smart. We found you also need to filter out "noise events" like VPN reconnections before the stats run. Otherwise a user connecting from a new coffee shop every week gets flagged unnecessarily. A quick pre-search filter for known noisy subnets cleaned that up.
Automate the boring stuff.
You're asking the right foundational question. The gap is almost always detection logic, but you can't even diagnose that until you prove the data pipeline is solid. Coming from data analysis, think of it as debugging a complex join that's returning null.
First, isolate a single manual finding. Get the exact source host, destination host, and timestamp. Don't search for parsed fields yet. Search the raw logs for that source host at that time and look for the actual Windows 4624 event text. Is it there? Good. Now, check if the key correlation atoms are present in the raw text: `TargetLogonId` and `TargetServerName`. If those are "-" or missing, your parsing is broken and no rule will work. Licensing can also cause this, as some log sources get filtered upstream.
If the data is intact, the most reliable technique I've used is a two-stage approach. Start with a simple frequency baseline for each source asset: count of unique destinations via logon type 3 over 7 days. A jump server will have a high count, a user workstation a low one. Then, pair that with a rule alerting on successful Type 3 logons from sources with a historically low count to sensitive destination server groups. This avoids the static list problem others mentioned.
Service accounts need entirely separate logic, treating them as entities with their own predictable patterns. A service account performing a Type 3 logon to a new server is often a bigger signal than a user doing it.
You're right about the grace period problem, but seeding from a template is just another static list with extra steps. The real cost is alert fatigue from false positives during that learning window.
I handle this with a temporary exclusion list fed by our IaC pipeline. When Terraform creates a management server with a specific tag, it drops the hostname into a lookup file for 48 hours. The scoring search checks that list and skips the baseline calculation altogether for those hosts. After the grace period expires, they fall into the normal scoring logic.
It's a band-aid, but cheaper than tuning down the entire rule sensitivity.
- elle
Welcome! Your data analysis background is going to be your secret weapon here, seriously. It's exactly the right way to frame the problem.
Everyone's already covered the crucial step of validating your raw log pipeline, which is 100% the first move. I'll jump straight to the detection logic part since that's where the fun is. For lateral movement, I've had the best luck with a two-part approach that's pretty simple to implement. First, you build a behavioral baseline for each account - not just "is this a service account," but what does normal activity for *this specific* service account look like over the last 30 days? Then, you look for deviations from that personal baseline, not just a global rule. A service account suddenly initiating an RDP session to a workstation it's never talked to before is a screaming signal, but your SIEM won't see it if you're only comparing against a list of "allowed" servers.
The second part is temporal correlation. An event at 3 AM might be meaningless on its own, but if it's part of a chain that started with a compromised user credential an hour earlier, that's your story. Tools like Splunk's transaction or stats commands can help stitch that session together, but you need those key fields like TargetLogonId to be present and parsed correctly. If they are, you can start building those session timelines automatically.
hugo
You've gotten some really solid advice, especially the step of validating your raw logs first. That's always step zero.
Assuming the data pipeline checks out, I've found one of the most effective (and simplest) techniques is to stop thinking about single events and focus on sessions. Look for the *chain* of a logon (4624) to a specific process creation (4688) using the TargetLogonId field. If you can tie a weird, after-hours RDP session to the instant launch of PowerShell or a remote service, that's your signal. The log sources are often all there, the correlation is just missing.
Have you looked at the user context for those "unexpected service account" activities? Sometimes the rule logic only flags admin accounts, but a standard user account doing something like scheduled task creation from a remote session is a massive red flag.
Automate everything.
Focusing on the data validation step is key. When you mention checking the raw event text for populated fields, is that something you do in a development or test Splunk instance first, or do you run those searches straight in production? I've been hesitant to do ad-hoc searches on the live logs.