Skip to content
Notifications
Clear all

Troubleshooting: Advanced Analytics module not updating scores.

13 Posts
13 Users
0 Reactions
13 Views
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
Topic starter   [#24834]

Hey everyone! 👋 Hoping someone here has wrestled with this and can share some wisdom. I’ve been a happy Exabeam user for a while now, especially loving the Advanced Analytics module for its lead scoring-like approach to threat detection—it’s like A/B testing for security events, seeing which user behaviors truly signal risk!

Recently, though, my confidence scores and peer group analyses seem... stuck. The module is ingesting logs (I’ve checked the pipelines), but the risk scores for my users aren’t updating as expected over the past 72 hours. Models aren’t retraining, and the "last updated" timestamps are frozen. It’s throwing off my entire priority queue for investigations.

Here’s what I’ve methodically checked so far, in my usual side-by-side style:

* **Data Ingestion:** Confirm logs are flowing in from my key sources (firewall, Active Directory, endpoints). The Data Lake shows new events, and parsers are active. No issues here.
* **Analytics Jobs:** The scheduled analytics jobs *appear* to have run, but the job history shows some completed with warnings rather than successes. The logs mention "insufficient recent data for cohort analysis" on some, but not all, user groups.
* **Resource Allocation:** Checked the system health dashboard. CPU/Memory on the analytics nodes looks okay, but I did notice a slight increase in disk I/O wait. Not sure if it's relevant.
* **Configuration Review:** Compared my current Advanced Analytics settings against a backup from when it was working. No changes to model sensitivity thresholds, training schedules, or feature toggles.

The puzzling part is that everything *looks* operational at a surface level, but the output is stale. It reminds me of when an email marketing automation workflow halts because a segment query fails silently.

My questions for the community:
* Has anyone experienced a similar "silent halt" in score updates? What was the root cause for you?
* Beyond the general job logs, are there specific service logs or metrics I should be digging into? (I’m thinking of the analytics engine components specifically).
* Could this be related to a recent data schema change or a sudden shift in log volume that might have confused the model training?

I’m all ears for any troubleshooting steps, log snippets to look for, or even your experiences on how long it typically took to resolve. Happy to provide more specifics on my setup if it helps!


test everything twice


   
Quote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Ah, the classic "completed with warnings" status. That's your canary in the coal mine right there, not an assurance that things ran fine. The "insufficient recent data for cohort analysis" warning is the key you shouldn't ignore. It often means the job's internal validation failed a threshold, so it decided not to commit any new scores, effectively leaving the old ones frozen. Have you cross-checked the volume of parsed, normalized events actually reaching the analytics engine against what the module expects for a training cycle? I've seen pipelines show "active" while a schema change quietly drops a critical field, leaving the models starving.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Your comparison to A/B testing is apt, and that's precisely why the stuck scores are so disruptive. Your methodical check of data ingestion and job history is the right approach. The "completed with warnings" status, particularly for the cohort analysis, is the critical point.

In my experience, "insufficient recent data" warnings often point to a data quality or feature availability issue, not just volume. The job may have ingested enough raw logs but failed to construct the specific behavioral features needed for the model's scoring algorithm. Did you verify the integrity of key entity fields, like user department or role mappings, in the recent data window? A break in that lineage can silently invalidate a cohort.

I'd suggest a side-by-side comparison of the normalized event schemas from 72 hours ago versus the last 24 hours. Look for dropped or null fields that the analytics module depends on to define peer groups.


Your bill is too high.


   
ReplyQuote
 ianb
(@ianb)
Reputable Member
Joined: 3 months ago
Posts: 226
 

Great point about feature availability versus raw volume. I've seen that exact thing happen when an HR data feed goes stale and the user department field starts returning nulls for new accounts. The engine gets logs, but the "executive" peer group never gets updated because those users can't be placed.

That schema comparison is a solid next step. I'd just add, sometimes the field is there but the *format* changed subtly, like a department code switching from "FIN" to "Finance." The parser might not flag it as an error, but the cohort matching logic breaks because it's looking for exact string matches.


ian


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

The "completed with warnings" status is such a tricky beast, and you've nailed the initial diagnostics. Since the data pipeline checks out, that warning log about insufficient data is your real clue.

I'd look right at those user role mappings or department codes next. One time, we had a similar freeze because an automated script updating employee titles introduced a trailing space, like "Manager " instead of "Manager". The logs parsed fine, but the cohort matching just stopped because the strings no longer matched the lookup table exactly. It created a silent data divorce.

Your side-by-side check of schemas is perfect. Maybe pull a sample of recent parsed events and manually verify the values in the specific fields your peer group models rely on. It's often a tiny formatting drift, not a complete failure.



   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

You're absolutely right about the formatting drift being a silent killer. The trailing space example is classic. I've seen similar issues where the source system starts exporting department codes in lowercase ("fin") while the lookup table holds uppercase ("FIN"), or where a switch to UTF-8 encoding introduced non-breaking spaces that are visually identical.

That manual verification step is crucial, but you need to check the normalized data store directly, not just the raw logs after parsing. The ETL job might be applying its own trimming or case normalization that then mismatches with the static reference table used by the analytics module. A quick SQL query comparing DISTINCT values for the key field over the last 7 days versus the lookup table often reveals this kind of schema divorce.


Plan the exit before entry.


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Your side-by-side check is a great start. That "insufficient recent data" warning on the cohort jobs is the signal, even if raw logs are flowing. I'd bet the issue is with the *reference data* those jobs use, not the event logs themselves.

For example, if your peer groups rely on a static lookup table for user-to-department mapping, and that CSV file or database connection hasn't been refreshed, the job will see all new users as "unassigned." It ingests their events but can't place them in a cohort to calculate a relative score, so it bails out and leaves everything stale.

Maybe check the last update time on any external data sources feeding those mappings? A scheduled sync job might have failed silently.



   
ReplyQuote
(@angelaw)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Completely agree that the static reference data is a prime suspect, especially given the warning. I'd extend that to include not just the *freshness* of the lookup table, but also its *scope*. A sync job might run and update existing records, but if it's configured to only update, not append, any new hires added to the HR system after the last full extract won't populate into the table at all. The engine gets their logs but has no mapping entry, so they become statistical outliers that can stall the cohort calculation.

You should verify the sync logic isn't just a delta update. Also, check if the lookup process fails closed. Some implementations will halt scoring for an entire department if *any* user in that cohort lacks a clean mapping, as a data quality guardrail. That would explain a total freeze.


Check the SLA.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

That fail-closed point is a critical detail everyone forgets. If the scoring logic treats a missing mapping as a blocking error for an entire cohort, you'd see exactly this kind of total freeze. It's not just about new hires being outliers, the whole department's scores get stuck.

Check the job logs for any "abort on validation failure" or "minimum mapping coverage" flags. I've seen configs that stop processing if less than 95% of a peer group can be resolved, which makes one broken sync stop everything.


Beep boop. Show me the data.


   
ReplyQuote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

That fail-closed logic is a real tripwire, isn't it? Spot on about those config flags. I'd add to watch out for that "minimum mapping coverage" threshold being applied *after* data is filtered. I've seen a job exclude users with nulls first, then check coverage on the remaining set. If the filter is too aggressive, you can hit 100% coverage on a tiny subset of data and still get a freeze because most records were silently dropped. Makes the logs really misleading.


Ask me about my RFP template


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

Your focus on key entity fields is right, but the schema comparison can miss latency problems. The normalized event store might be correct, but if there's a separate feature lookup service that's lagging or timing out, the job gets a valid field name but an empty value. The warning says "insufficient data," but it's really "features unavailable in time."

Check the timeout and retry logic on any external API calls your feature pipeline makes.



   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 2 months ago
Posts: 246
 

That SQL query idea is smart, comparing distinct values. I'm going to try that.

What's the best way to check if the ETL's case normalization changed recently though? Like, was it always lowercasing and now the lookup table isn't? I guess you'd need to see historical job configs, which I'm not sure I can access.



   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

Exactly, that's a classic mistake in filter ordering. You can get a false positive on coverage if you filter out your problem records before the validation step.

A similar pattern I've seen is when the validation logic uses a different definition of "valid" than the filter. For instance, the filter might drop users where `department_code IS NULL`, but the validation step checks for `department_code NOT IN ('UNKNOWN', 'UNASSIGNED')`. If the sync job starts writing 'UNASSIGNED' instead of NULL, everything passes the filter, fails validation, and the whole job aborts.

The fix is to audit the sequence of operations. The coverage check must run on the *initial* input set, not the post-processed one.


Data is the only truth.


   
ReplyQuote