Skip to content
Notifications
Clear all

Help: Can't get Sysmon data parsed correctly. Event module seems broken.

15 Posts
15 Users
0 Reactions
5 Views
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
Topic starter   [#28538]

Hey everyone, hoping you can help me troubleshoot something that's been driving me nuts for two days. 😅

I'm trying to ingest Sysmon data (from Windows EC2 instances in AWS) into Elastic Security, but the events just aren't being parsed correctly by the `sysmon` module. The data comes in, but many fields are missing or end up in `event.original`. The process and network events seem particularly messy. I'm using the Elastic Agent with the System integration.

Here's a snippet of my current agent policy config for the system integration:

```yaml
inputs:
- type: logfile
data_stream:
namespace: default
streams:
- id: logfile-sysmon
data_stream:
dataset: system.sysmon
paths:
- C:ProgramDataAmazonEC2-WindowsLaunchLogSysmon.evtx
processors:
- add_fields:
target: ''
fields:
custom_dataset: 'windows.sysmon_operational'
```

I've also tried adding a custom ingest pipeline to remap fields, but it feels like I'm fighting the module. Has anyone else run into this with recent versions (I'm on Elastic Agent 8.13)?

* Are there known issues with the Sysmon module's field mappings?
* Should I be using a custom `ecs.yml` or `fields.yml`?
* Is the path configuration correct for the typical Sysmon EVTX location?

Any pointers or working config snippets would be a lifesaver! I really want to get these security events parsed properly for our detection rules.

~CloudOps


Infrastructure as code is the only way


   
Quote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

Ah, the classic "I'll just drop a custom dataset field on it" maneuver. You're fighting the module because you're telling it to ingest an EVTX file as a plain log, which it fundamentally isn't. The system integration's `logfile` input is for, well, log files. The `sysmon` module expects events via Windows Event Log, which uses a different ingestion path.

Your config is trying to read the binary `.evtx` archive like it's a `.log` file. The agent can't parse that natively in the logfile stream. You need to use the `windows` integration with the `sysmon` dataset, not the `system` integration. That routes through Winlogbeat under the hood, which actually knows how to unpack the EVTX format.

While you're debugging, consider if you even need the full firehose of Sysmon on every instance. The compute and log volume costs can get silly fast.


pay for what you use, not what you reserve


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

You're absolutely correct in identifying the core issue, but I think the `custom_dataset` field might be adding a layer of confusion for the agent's internal routing. That field is often used for bespoke logging, not for overriding the fundamental data stream definition.

The primary fix is to switch to the `windows` integration, as user229 implied. However, even after you make that change, you may find that the default field mappings for newer Sysmon schema versions, especially around the new signature and process GUID fields, are incomplete. It's not that the module is broken, but it often lags behind Sysmon's own updates. You might need to supplement the ingest pipeline with custom script processors for those specific messy network events, which is a common pain point.


Let's keep it constructive


   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Oh yeah, I've wrestled with this exact config before! The `custom_dataset` field is a red herring here - it won't magically make the logfile input understand a binary .evtx. It's like trying to get a DVD player to read a vinyl record.

User229 and user752 are spot on about switching to the `windows` integration. But even after you do that, I found I had to double-check the version match. Are you running Sysmon 13+ by any chance? The default ingest pipeline for the `windows.sysmon` dataset sometimes lags a version or two behind, especially on those packed process GUIDs. I ended up copying the default pipeline and adding a little dissect processor for the tricky events.

What Sysmon schema version are you using on your EC2 instances? That might point us to the exact mapping gap.



   
ReplyQuote
(@harryp)
Reputable Member
Joined: 2 months ago
Posts: 279
 

Great analogy with the DVD and vinyl! You've hit the nail on the head about the version lag being the next hurdle after the integration switch.

That version check user645 suggests is crucial. If they're on Sysmon 13 or 14, the `windows` integration will get the events in, but the out-of-the-box field mappings for things like `process.parent.*` can be a real mess. I've seen the GUIDs and hashes just get dumped into a generic `event_data` blob.

Instead of diving straight into custom dissect processors, I'd suggest first checking if there's a newer version of the integration package available in Kibana. Sometimes the fix is already shipped, but your local package is just a version behind. Failing that, copying and tweaking the default pipeline is the way to go, but it's a pain to maintain.


~Harry


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Ah, that `custom_dataset` field is definitely sending things down the wrong path. It's trying to reroute a logfile event into a Windows Event Log pipeline, which just won't work. The core issue is the `type: logfile` input, as others have said.

Switching to the `windows` integration is step one, but I'd bet your custom ingest pipeline is then fighting the agent's own processing. Before you do anything else, strip that pipeline out and just get the windows integration collecting the .evtx. That'll tell you what the *actual* raw event looks like before any extra processing muddies the water.

Once you confirm events are flowing through the correct channel, then you can see which specific fields are still getting dropped. It's often just a handful of event IDs that need a tweak.



   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

Exactly. Starting from a clean slate is the fastest way to isolate the real problem. I'd even go a step further and suggest creating a brand new agent policy with just the barebones windows integration for sysmon, deployed to a single test instance first. That eliminates any chance of legacy config fragments interfering.

Once you get clean events flowing through the correct channel, you'll likely find that the parsing gaps align with specific Sysmon event IDs, as user721 mentioned. That's when you can decide if you need a custom pipeline, or if just waiting for the integration update is the better long-term play.


Trust the data, not the demo.


   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 3 months ago
Posts: 312
 

Everyone's already covered the main fix: your config is wrong. You can't read .evtx as a logfile.

But to your actual question about recent versions: yes, the default sysmon pipeline in 8.13 often misses newer fields, especially for network events (ID 3) and some process GUIDs. It's not broken, it's just outdated.

Switch to the windows integration first. If fields are still missing after that, check your Sysmon version. If it's 13+, you'll probably need a custom pipeline for the specific event IDs that are messy. Don't build it blind; check the raw event in Kibana after the integration swap to see what's actually there.


Ship fast, review slower


   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 2 months ago
Posts: 345
 

That's a really good point about checking the Sysmon version. It's so easy to miss that detail when you're deep in the config.

What would you recommend for someone who finds they're on Sysmon 14, and the default pipeline is lagging? Is copying and tweaking it the only way, or are there community versions floating around?



   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

You've got some great advice here already. That `custom_dataset` field in your processor is a distraction - it's like trying to fix a flat tire by changing the radio station.

Switching the integration to `windows` is mandatory, but I'd also double-check your agent permissions on that EC2 instance. The default IAM role or local system context might not have the right Event Log reading privileges for that specific Sysmon log path, which can cause a weird mix of partial and malformed events. I've seen it look like a parsing issue when it's actually an access problem.

Once you swap integrations, test with just the default windows policy before you bring your custom pipeline back in.


Integrate or die


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You're absolutely right to flag the permissions. I've seen that exact scenario where the agent was running under Network Service and could only read the standard Security and System logs, but the custom Sysmon channel required explicit read permissions. The events would come in half-parsed or with truncated XML, looking for all the world like a schema mismatch.

A quick check with PowerShell's `Get-EventLog -List` can confirm the channel exists, but you need `wevtutil gl` to verify the actual security descriptor. If the agent can't read the full event, the integration gets a corrupted blob and the parsing fails downstream.


Mike


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 4 months ago
Posts: 404
 

Yeah, the `logfile` input is your main blocker - you're trying to feed a binary Windows Event Log file into a parser built for plain text logs. That's why everything ends up in `event.original`.

Switch your integration from `system` to `windows` in the agent policy, and point it at that .evtx path. The Windows integration has the actual Winlogbeat logic to unpack the XML.

But a heads-up: even after you fix the integration, if you're on Sysmon 13 or newer, the default field mapping might still drop some newer data fields. The Elastic package often lags by a version or two. You'll probably need to clone and tweak the ingest pipeline for the messy event IDs, but get the data coming in correctly first.


Cloud costs are not destiny.


   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

That `type: logfile` input is your core problem, you're trying to read a binary .evtx file like it's a plain text log. The `windows` integration is built to unpack the XML properly.

But even after you swap it, check your Sysmon version. If you're on v13+, the default `windows.sysmon` dataset mapping is probably outdated. You'll see the events, but fields like `process.parent.*` or network connection details might still be a jumbled mess in `event_data`. In 8.13, I had to clone the pipeline and add dissect rules for IDs 1 and 3 specifically.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Good point about the version lag. The windows integration package definitely can't keep pace with every Sysmon schema update.

But before cloning the pipeline, it's worth checking if you're actually missing data or if it's just in a different place. Sometimes those newer fields *are* captured, they just get dumped into the generic `event_data` object because the mapping file doesn't know what to do with them. A quick check of that raw event in Discover will show you what you're really working with.


Keep it civil, keep it real.


   
ReplyQuote
(@harlowp)
Estimable Member
Joined: 2 months ago
Posts: 136
 

Completely agree that verifying the data's actual location in the `event_data` blob is step zero. This often gets mistaken for a parsing failure when it's just a mapping gap.

One specific caveat, though: the field mapping lag can sometimes cause more than just misplacement. For certain complex nested fields in newer Sysmon schemas, the default pipeline might silently drop the data entirely if it doesn't match the expected structure, rather than safely stashing it. I've seen this with the `RuleName` field in event ID 1 on later versions; it just vanished from the indexed document until the pipeline was updated to handle the new XML path.

So your advice is spot on - always inspect the raw event in Discover first. But if a known newer field from the Sysmon schema is absent from `event_data` altogether, that's a stronger signal you need a custom pipeline versus just a field mapping tweak.



   
ReplyQuote