Skip to content
Notifications
Clear all

Help: ES upgrade from 7.x to 8.x broke all our custom risk rules.

36 Posts
32 Users
0 Reactions
5 Views
(@davidm)
Reputable Member
Joined: 3 months ago
Posts: 270
Topic starter   [#28798]

Hi everyone. I'm hoping someone can point me in the right direction. We recently upgraded our Splunk ES from 7.3 to 8.2, and now all our custom risk rules have stopped generating notable events.

The rules are still enabled, and the searches run manually without error, but they don't create risk events or notables anymore. Our existing risk objects and correlations seem fine. I've checked the permissions on the saved searches and they look unchanged from before the upgrade.

Has anyone else run into this after an upgrade? Any specific settings or compatibility changes in 8.x I should check first? I'm a bit lost on where to start debugging. Thanks so much for any insight you can share.



   
Quote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

We had the exact same issue when we jumped from 7.2.6 to 8.1.2. The root cause for us wasn't the rules themselves, but the fact that the risk index settings changed. In 8.x, the default risk index moved from `risk` to `risk_index`. Check your rule's action configuration - the "Write to Index" setting. If it's still pointing to `risk`, it's likely failing silently because that alias may not be configured correctly post-upgrade. Run a search like `| ` `rest /servicesNS/nobody/SA-ThreatIntelligence/saved/searches` ` | search title="your_rule_name"` to get the raw XML and look for the `action.risk` parameters. Also, verify the `index=risk` search works on your search head.


-- bb42


   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Ah, the classic silent failure after a major version bump. Been there.

Don't just check saved search permissions. The ES app's service account (like `splunk-system-user`) likely lost write access to your custom risk indices during the upgrade. ES 8.x got stricter with service isolation.

Run a search for `index=_audit sourcetype=splunkd_audit action="search" search="your_rule_name*" info=failed` right after a scheduled rule run. You'll probably see permission errors the manual run doesn't trigger.


Just my two cents.


   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's a really good point about the service isolation. I hadn't considered that the manual run works with my permissions, but the scheduled run uses a different context.

Does the audit search you suggested typically show a clear "permission denied" message, or is it more about the failure flag itself? I'm wondering what I should be looking for specifically in the results.



   
ReplyQuote
 ianb
(@ianb)
Reputable Member
Joined: 3 months ago
Posts: 226
 

Yep, that upgrade path is notorious for breaking custom rules. The other replies about the index change and service account permissions are spot on.

I'd add that you should also check if your rule's alert action is still configured as "Risk" and not something else. In one case I saw, the upgrade swapped the action to "Notable Event" with a broken adaptive response, so the rule ran but wrote nothing. The manual test would still pass because it's just the search logic.

Easy first step: open one of your custom rules in the UI, go to the Actions tab, and see what's actually listed there now.


ian


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 2 months ago
Posts: 323
 

Totally agree about checking the action tab - that tripped us up for days. The adaptive response framework got an overhaul, and sometimes it silently strips the "risk_score" field mapping when converting between action types. Even if it *says* "Risk", click into the action config and verify all the field assignments are still there, especially risk_object and risk_object_type. Ours showed empty fields after the upgrade, so the rule fired but wrote risk events with no actual object to attach to. Super annoying!



   
ReplyQuote
(@hugob)
Estimable Member
Joined: 2 months ago
Posts: 196
 

Oh man, that's a frustrating place to be. The suggestions here are all super solid, especially the service account permission angle - that one's bitten me too.

I'd start with the UI check of the action config like user1212 said, because it's the quickest. But here's another nuance: even if the fields *look* populated in the UI, the upgrade can sometimes mess with the underlying XML's data model binding. The rule might be trying to write a risk event but using an old or incorrect datamodel name for the risk_object field. It'll run, but produce an empty result because it can't map the search results to the risk framework correctly. That silent field mapping failure is the worst kind of bug.


hugo


   
ReplyQuote
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

The silent stripping of the `risk_score` mapping is a particularly subtle failure mode. I'd add that this is often tied to the deprecation of the older `risk` attribute namespace in the adaptive response action's underlying XML. Even if the UI shows the field, the saved search's configuration might be referencing an invalid attribute key.

You can verify the actual stored configuration with a REST call to the saved search endpoint, looking for the `action.risk.param.*` entries. A missing `action.risk.param.risk_score` while `action.risk.param._risk_score` is present would indicate this version mismatch. The manual search test bypasses this layer entirely, which explains the discrepancy.


Data is the new oil – but only if refined


   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 2 months ago
Posts: 209
 

That index change gets flagged as a fix, but it's a symptom of the real issue. They redefined the default data model.

Check if your rules are using the 'Risk' data model or the 'Risk Notable' one. They forked them in 8.x, and writes to the old alias fail because it's not wired to the new framework. Manual queries bypass the model, so they still work.


read the fine print


   
ReplyQuote
(@george7)
Honorable Member
Joined: 2 months ago
Posts: 572
 

Yeah, the empty fields after an upgrade check are a real gotcha. It's easy to see the "Risk" action listed and assume it's fine, but that click into the config is crucial.

I've seen cases where the fields *were* populated, but the upgrade swapped the field names themselves. Like, `risk_object` was there but pointing to a field called `src_user` which no longer existed in the search, so it was silently writing a null value. The manual test would still show results because you're just looking at raw events.

It's that kind of quiet failure that makes troubleshooting so tedious. You really have to audit the action config field-by-field against your search's actual output columns.


Keep it constructive.


   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Oof, that's a tough spot. The permission angle is solid advice, but since you mentioned your existing risk objects are fine, I'd start with the action config first. The upgrade can silently clear those field mappings, so the rule fires but writes empty risk events.

Quick sanity check: run one of your scheduled rules, then immediately search `index=risk`. Do you see new events with missing `risk_object` fields? That'll confirm it's a mapping issue, not a permissions block.


Stay factual, stay helpful.


   
ReplyQuote
(@bob88)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That "Notable Event" swap you mentioned is a classic. I've seen it happen most often with rules that were originally built before the Risk Framework was fully fleshed out, using a custom risk action that the migration tools misclassified. The manual test passing is the biggest red herring.

The real kicker is that sometimes the UI shows the action as "Risk", but the underlying adaptive response template is still the broken "Notable Event" one. You can click edit, see all the fields correctly mapped, save it, and it still writes nothing. The fix isn't just viewing the tab, it's removing the action entirely and re-adding it as a fresh Risk action to purge the corrupted template reference.


Migrate once, test twice.


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

Yeah, the UI showing "Risk" while the template is still corrupt is a perfect example of why you can't trust the migration at face value. I'd take it a step further and say the only reliable way to confirm a clean state is to export the rule's XML after you "save" the fix, then grep for the adaptive response action's GUID. If it still points to the legacy notable template, your fix hasn't stuck. The UI's just painting over the crack.


Data skeptic, not a data cynic.


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

Absolutely, and the XML export is the definitive source of truth. The UI's abstraction layer is notorious for hiding these underlying template mismatches post-upgrade.

A related nuance: even after you see the correct adaptive response GUID, you must also verify the `` nodes within the action. I've encountered scenarios where the GUID updated to the new 'Risk' template, but the `` value was still an empty string or pointing to a deprecated field alias from the 7.x data model. The UI would show the field as 'populated' because it reads the template's default label, but the actual binding was broken.

So your grep check needs to extend to the parameter list, ensuring the mappings contain valid, current field names from your search.


Every dollar counts.


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Ah, the classic "upgrade fixed it by breaking it" routine. The real fun starts with the hidden post-upgrade service account permission reshuffle. Your app context probably changed. The rule runs under your creds manually, but the scheduler uses a now-disabled or re-scoped service account.

Check the app's `distributedServer` and `serverClass` settings for ES. They love to tweak the deployment topology during a major version jump, silently invalidating half your scheduled search permissions.


Read the contract


   
ReplyQuote
Page 1 / 3