Skip to content
Notifications
Clear all

How do I get started writing my own correlation searches? Any good templates?

8 Posts
8 Users
0 Reactions
34 Views
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
Topic starter   [#22356]

Hey everyone. I've been using Splunk ES for a few months now, mostly relying on the built-in correlation searches and security content. It's great, but I'm hitting a wall where I need to start building custom searches for some of our specific internal workflows and oddball SaaS apps we use.

Coming from a CRM/RevOps background, I'm used to building alerts in Salesforce or HubSpot for weird data patterns, so the logic makes sense, but the SPL and the ES framework feel like a different beast. I get a bit lost on the "right" way to structure things within the ES app itself.

So, for those of you who've built your own, how did you start? Are there any good templates or a basic structure you follow? I'm curious about things like:
- How you set the data model alignment (do you always use datamodel=...?)
- Best practices for naming and severity assignment.
- Any good examples of a simple, yet effective, custom correlation search you built from scratch.

I'm not looking for anything super complex to startβ€”maybe something like detecting a user's successful login from two geographically impossible locations within a short timeframe, but for a non-AD application. Just trying to understand the scaffolding. 😅

Any migration stories from generic Splunk alerts to proper ES correlation searches would be super helpful too.



   
Quote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

Start with the built-in ones. Go to Configure, then Correlation Searches in ES and open a few in edit mode. That's your template. Copy one that's close to what you need and modify the SPL.

For your questions: you don't always need `datamodel=`. If your SaaS data is already CIM-compliant, use the tstats command on the model. If it's weird log data, just search the index/sourcetype directly and map fields later. Naming should follow the ES convention: "Application - Descriptive Name - Rule". Use the built-in severity matrix as a guide, don't just pick high because it feels right.

Your geographic impossible login example is a good start. The structure is already in the ES content. Look at "Access - Multiple Successful Logins From Geographically Distant Locations - Rule". Swap the identity and authentication data model references for your specific index and user/location fields. Test it against a small timeframe first with a `| stats count` to verify logic before you schedule it.


Where is your SOC 2?


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Building from user323's point about CIM compliance, I'd add a practical distinction. If your SaaS app logs are already normalized via Technology Add-ons or Common Information Model (CIM) fields, absolutely use `datamodel=` and `tstats`. The performance benefit is massive for scheduled searches. If they're not, starting with a raw search on the index is fine, but you should immediately map the key fields (`user`, `src_ip`, `action`) to their CIM equivalents using `eval` or field aliasing in your `props.conf`. This future-proofs the search and lets you integrate with other ES features later.

For your impossible travel example on a non-AD app, the core SPL structure is straightforward. You'd start with the raw logs, filter for successful logins, then use `transaction` or `stats` to group by user with a time window. The geographic lookup would typically be done with a `lookup` to a geolocation table based on `src_ip`. The main deviation from the built-in search is just your source data.

One template I frequently reuse is this pattern for anomaly detection on low-volume events:
```
| tstats summariesonly=true count from datamodel=Authentication where Authentication.action="success" by Authentication.user,_time span=1h
| streamstats window=5 current=true avg(count) as avg_count stdev(count) as stdev_count by Authentication.user
| eval threshold = avg_count + (3 * stdev_count)
| where count > threshold AND count > 5
```
This flags users whose login count in the last hour is more than three standard deviations above their rolling average. You swap out the data model, the action field, and the grouping key. It's simple, statistically grounded, and avoids static thresholds.


data is the product


   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Great point about mapping fields early. I've seen too many teams skip that step and then struggle to adopt new ES use cases later. That props.conf work is a bit of upfront pain, but it pays off.

Your anomaly pattern is solid for low-volume stuff. One thing I'd add - when working with raw SaaS logs that aren't CIM-ready, test the search with `eventstats` instead of `transaction` first. It can be way more forgiving with weird timestamp formats, which these apps love to throw at you.


Trust the trial period.


   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

Good call on mapping fields early. The props.conf step is tedious but it's a one-time cost. If you don't do it, you'll be rewriting every search later when you want to use Risk-Based Alerting or asset/identity correlation.

One caveat to the tstats performance advice: it only works if your data model accelerations are actually built. For a brand new, low-volume data source, a raw search might run faster initially until the summaries populate.



   
ReplyQuote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

I've been in this exact spot, trying to translate workflow logic into ES. Since you mentioned coming from CRM alerts, the mapping might be less daunting if you think of building the search first, then retrofitting it into the ES framework. For your impossible travel example on a non-AD app, I'd do just that.

Start by writing the SPL that works on the raw index to find your pattern, making sure it runs and returns the right results. That's the hard part, honestly. Only after that works would I worry about CIM mapping or slotting it into a correlation search template. It helps to separate the logic-building from the ES-compliance stress.

On severity, I'm still figuring that out myself. I found the ES matrix a bit abstract initially. Lately I've been using the built-in searches as a reference library: if I find one that feels like a similar impact level for my org, I just mirror its severity and confidence settings. It's not perfect, but it gets you past the initial guesswork.



   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Copying an existing search is the right start, but don't just tweak the SPL. Open the search manager and look at the whole configuration - the schedule, the action triggers, the drill-down. That's the framework.

Ignoring the CIM mapping for a raw index works initially, but you'll regret it. Do the props.conf work immediately. It's not optional if you ever plan to use risk scoring or identity linking.

For your impossible travel example, the built-in ones are over-engineered for a simple app. Build the raw search first:

```
index=saas_logs action=login success=true
| bucket _time span=1h
| stats earliest(_time) as first_time, latest(_time) as last_time, values(country) as countries by user
| where mvcount(countries) > 1
```

Test that. *Then* wrap it in the ES template with your mapped fields. Naming: keep it simple. "SaaS App - Impossible Travel - Rule". Severity should mirror an existing one with similar business impact, not your gut feeling. Look at the actual incident response load it will create.



   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Everyone overcomplicates this. You don't need a "template," you need a working search.

> how do you set the data model alignment
You don't, at first. Forget the ES app. Write the SPL that finds your pattern in the raw logs. If it works, then you can waste time trying to cram it into a datamodel later.

The impossible travel example is the classic trap. The built-in search is a mess of macros and datamodel calls. For your oddball SaaS app, just find the logins, bucket them, and check for two different locations. If that logic works, you've done the hard part.

Naming and severity are just bureaucracy. Copy whatever format they use in the existing list. Severity is a guess until something actually happens.


your mileage will vary


   
ReplyQuote