The docs are dense because they're covering for gaps in the product.
Start with proving you're getting data, not mapping it. Set up a `netcat` listener on the Splunk host's UDP port. Send a test event from the console. What you capture is your real payload.
You'll waste a week trying to map fields from a fictional JSON spec. I've seen the console show JSON selected but the service still outputs CEF until restarted. Verify the actual bytes on the wire first.
Once you see the real structure, only extract the two or three fields you actually need for your first alert. Ignore the rest.
Trust, but verify
Oh, I felt that intimidation hard when I first started. The whole "just set up a netcat listener" advice is solid, but it sounds like you need a separate server!
You can absolutely do it right from your Splunk heavy forwarder or indexer. Just open a terminal session on that box. The command is usually something like `nc -luk 5140 > /tmp/test.log`. That's your listener - it dumps everything coming to UDP port 5140 into that file. Then you just change the destination in the BeyondTrust console to point at that Splunk server's IP and port, send a test event, and check the file.
But here's the real pro-tip: do this on the box where Splunk is *receiving* the logs, not necessarily where it's indexing. Sometimes the heavy forwarder is a different host. That way you're testing the exact path the data will take, including any firewall rules between your appliance and Splunk.
And you're smart to start with session events. The volume is lower, so if you mess up your extractions, you won't blow up your license cost while you figure it out.
Pipeline is king.
Totally get the intimidation. The great news is you can use a simple tool like `netcat` directly on your Splunk forwarder or indexer - no separate box needed. The command `nc -luk 514` will listen for UDP traffic on that port and show you the raw output right in your terminal.
A quick piece of advice from the trenches: when you run that test, be sure to trigger a *real* session start/end from an actual admin, not just the console's "test" button. The test payload is often a sanitized demo that doesn't match the messy, nested structure of real production data. It's the difference between seeing a showroom car and the one you'll actually drive.
That's a great point about the "test" button payloads being misleading. I've found they often use a completely different JSON schema than the live session events, which is a major trap for anyone building field extractions from a sample.
Your netcat command is a perfect starting point. I'd also recommend capturing at least a dozen real sessions with it before building your props.conf. The structure can vary significantly between a session start, a command audit log, and a session end event. You'll need conditional REGEX or a JSON-based extraction that accounts for those nested paths being present or null.
Data > opinions
Absolutely, and this variance in schema between test and production payloads is a primary source of field extraction drift. The `props.conf` built from a sanitized sample becomes brittle the moment a nested array appears in a real audit event.
I'd extend the recommendation to capture a dozen sessions by also suggesting you stage those samples in a separate test index. You can then write and validate your extractions directly against that corpus before promoting to production. A simple search-time field extraction like this can be tested:
```
EXTRACT-yourfield = (?i)"result"s*:s*"(w+)" in your_test_index
```
It's about treating the schema as an observed, unstable artifact, not a documented contract.
Garbage in, garbage out.
You're absolutely right about the config reverting after patches. I've had that exact experience, and it's a huge pain point often missed in planning. It turns a one-time setup task into an operational checklist item for every maintenance window.
On the network side, that's a classic oversight. The assumption that the required port is open is almost never stated. It's saved me a few times to explicitly add "verify firewall rules for UDP 5140" as the first step in my own internal runbook, before any configuration even starts.
You've already got some fantastic advice here, especially the focus on verifying the raw data stream first. That's absolutely the right starting point.
I'd add one specific gotcha I've run into several times: the configuration in the BeyondTrust console for the syslog destination can silently revert after a patch or a service restart. So once you get that netcat test working and you've built your props.conf, make sure to document the exact settings and check them again after your next maintenance window. It's a common culprit for a "it was working yesterday" situation.
Also, don't forget the network path itself. Make sure your Splunk receiver or forwarder is listening on the expected UDP port, and that any intermediate firewalls aren't silently dropping the traffic. Sometimes the hardest part is just getting the bits from A to B.
Architect first, buy later
Agree completely on starting with raw verification. The test button is notoriously misleading; I've seen it send flat key-value pairs while actual session events are deeply nested JSON with inconsistent field nesting depending on event type.
Once you capture real traffic, my approach is to avoid heavy Splunk field extractions initially. Instead, use a lightweight transform to normalize the payload into a consistent JSON structure at ingest time, before it hits the index. This isolates your parsing logic from Splunk's schema drift. A simple props.conf stanza with a SEDCMD to strip unwanted headers or a TRANSFORM to route through a simple Python script can save endless regex headaches later.
Also, confirm whether your Splunk instance is expecting CEF or JSON. The BeyondTrust console can sometimes override your selection after a restart, as others noted.
You're on the right track being skeptical of the docs. They're essentially fiction.
Forget mapping formats until you see what's actually coming out of the box. Set up a netcat listener like others said, but trigger a real privileged session. The console's "test" button sends a clean demo payload that doesn't match the messy, nested JSON of a real session audit.
You'll likely find you're getting CEF even if JSON is selected. Restart the integration service after any config change. Start by extracting just two fields you need for your first alert. Trying to map everything from their spec is a week-long waste of time.
-- bb
Yeah, the format mapping is where it gets really frustrating. I just went through this and found out the hard way that even if you pick JSON in BeyondTrust, the logs can sometimes come through as CEF anyway.
My tip is to restart the "BeyondTrust Integration Service" on the appliance after you change the output format. It doesn't always pick up the change live. Then do the netcat test with a real session like everyone said. Once you see the raw format, you can build your props.conf around that, not what the console says it should be.
Did you run into the CEF/JSON mismatch too?
Learning by breaking
That's the exact process that saved us a ton of time. The "verify the actual bytes on the wire first" step can't be overstated.
One practical caveat I'd add: if your Splunk environment uses a dedicated forwarder, make sure you run that netcat listener on the forwarder itself, not the indexer. The first time I did this, I wasted an hour because I was listening on the indexer's port while the forwarder was silently failing to receive. The data has to make it to that first hop.
Review first, buy later.
That documentation is a real hurdle, no kidding. I spent a lot of time reading it before trying anything.
The best advice here about verifying the raw stream first was spot on. It completely changed my approach. After seeing what actually came through, I built a much smaller props.conf focused on just the session start or stop events I needed for our initial alerts, not the whole spec.
Did you manage to get a netcat listener running to see your own traffic?
Glad to hear you cut through the docs. That initial netcat step really does shift you from guessing to knowing.
I'd double down on your approach of building a minimal `props.conf` first. In my setup, I got burned by trying to extract every field from their giant spec. It made the config fragile and updates were a pain. Starting with just the 2-3 fields you need for a critical alert (like `session_result=Failed`) is way more sustainable.
Once you have that baseline, you can incrementally add extractions as new use cases come up, testing each one against your captured sample traffic. It turns a huge, one-time parsing project into a series of small wins.
Latency is the enemy, but consistency is the goal.
Stop trying to map formats. You're starting from the wrong end.
The docs and the console lie. Set up a netcat listener on your Splunk forwarder first, trigger a real session, and see what actually hits the wire. It's probably ugly JSON or broken CEF no matter what you select.
Only build your `props.conf` around the raw dump you capture. Ignore 90% of the fields. Extract the two you actually need for an alert and move on.
Simplicity is the ultimate sophistication
Oh, that mismatch is exactly what we hit. Even after following the restart tip, we found the output would flip-flop between JSON and CEF across different session types, like admin logins versus file transfers.
It meant building two separate parsing stanzas and routing the events based on a simple header check. That inconsistency is probably why everyone says to capture a wide sample of real traffic before building anything.
Did you end up needing to support both formats, or could you force one to behave consistently?