Skip to content
Notifications
Clear all

Just automated risk score adjustments based on our ticket system.

22 Posts
22 Users
0 Reactions
119 Views
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
Topic starter   [#22208]

Having recently completed an integration between Exabeam's New-School SIEM and our internal IT service management platform, I believe there is significant, under-discussed value in leveraging external business context to modulate security risk scores. The canonical use case—boosting the risk of a user who just submitted a resignation letter—is merely the entry point. Our implementation focuses on a more nuanced, bidirectional adjustment model based on ticket state, category, and resolution.

The core premise is straightforward: a security alert generated by Exabeam carries a certain risk score. However, if that alert correlates to an open ticket where the activity has been deemed legitimate (e.g., a "server patching" ticket justifying unusual after-hours logins from an admin), the risk should be attenuated. Conversely, the absence of a justifying ticket for anomalous behavior should provide a positive weighting factor. The technical execution, however, involves several distributed systems challenges.

Our architecture hinges on a lightweight middleware service, acting as an enrichment plugin within Exabeam's ecosystem. This service listens for notable asset or user sessions from the Exabeam Data Lake, performs a near-real-time query against the ITSM's REST API, and applies adjustment logic before the session is evaluated by the Behavioral Analytics engine. The critical design points were:

* **Idempotency and Retry Logic:** ITSM API calls can fail. Our enrichment service must store its adjustment decisions in a local, persistent cache (we used Redis) to ensure repeated processing of the same session doesn't lead to score drift.
* **Stateful Correlation:** Mapping a security event to a ticket is not trivial. We key on a combination of `asset_hostname` + `time_window` for infrastructure-related alerts and `username` + `activity_type` for user-centric ones.
* **Adjustment Rules as Code:** We avoid GUI-based rules for maintainability. The adjustment policies are defined in a declarative YAML configuration, version-controlled and deployed with the service.

```yaml
# Example adjustment policy (excerpt)
adjustment_policies:
- name: "legitimate_change_window"
trigger_event_category: "Authentication"
correlation_key: "asset_hostname"
itsm_ticket_query: "category='Change Management' AND status IN ('In Progress', 'Implemented') AND affected_assets CONTAINS {{asset}}"
score_adjustment: -30
max_adjustment_cap: -50
conditions:
- field: "event_operation"
operator: "IN"
values: ["Success", "Failure"]
```

We observed a measurable reduction in false positive notable user sessions—approximately 18% over a 90-day evaluation period—specifically in change management and incident response scenarios. However, pitfalls emerged:

* **Latency Introduced:** The synchronous API call to the ITSM adds 200-300ms to session processing. We had to tune the Exabeam session watermark accordingly to avoid pipeline stalls.
* **Data Model Discrepancies:** Aligning Exabeam's asset inventory (often from Active Directory) with the CMDB in the ITSM required a normalization layer, which became a significant sub-project.
* **Adjustment Overreach:** An initial rule that overly aggressively reduced scores for any open ticket led to a missed true positive where an attacker had exploited a known, ticketed vulnerability. We introduced a `max_adjustment_cap` and a blacklist of high-severity alert types that cannot be fully neutralized.

The integration ultimately shifts the SOC's workflow. Analysts now see a contextual tag on the notable session timeline: "Correlated with ITSM Ticket CHG-12345 (Change Management)". This provides immediate investigative context without requiring a manual tab switch. The more ambitious, and still ongoing, extension is feeding security incident tickets created in Exabeam back into the ITSM, creating a closed-loop system where risk scores influence operational priority, and operational context refines risk scores. I am interested in hearing from others who have attempted similar contextual integrations, particularly regarding how you handle the eventual consistency problem between the SIEM's data lake and the external system of record.



   
Quote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That bidirectional piece is critical. Without it you'd just be suppressing alerts.

But how are you mapping ticket categories to risk adjustments? We tried something similar and ended up with a huge lookup table because "server patching" in one team's system was "infrastructure change" in another.

Did you have to normalize all your ITSM data first?


—cp


   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

You've hit the nail on the head with the normalization problem. We avoided the giant lookup table by taking a simpler, more surgical approach. Instead of trying to map every category, we only mapped a handful of high-fidelity ticket types that had a clear, predefined security implication.

We worked with the service desk to create three standardized ticket categories specifically for this integration: "Planned Elevated Access," "Authorized Data Transfer," and "Security Investigation - Legitimate." Only tickets in these buckets trigger adjustments. Everything else is ignored. It means we're not using the full breadth of ticket data, but the signal we do get is clean and reliable.

The trade-off is coverage, but it eliminated the mapping maintenance nightmare. We found most of the value came from that small set of predictable, high-trust scenarios anyway.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

We went the opposite direction and automated the mapping. Our integration scrapes the ITSM API, pulls all active category/priority combos, and dumps them into a small service that uses fuzzy matching and a manual override config file.

It's more up-front work but now when the networking team decides to rename "router config" to "network device provisioning", we don't have a broken integration. The service suggests a new mapping, we approve or reject it. It's a one-minute task for an on-call instead of a project.

Your approach is cleaner for sure, but I'm too paranoid about missing a critical ticket because someone used the "wrong" category in a real incident.


Automate everything. Twice.


   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

The normalization challenge is the real beast here, isn't it? You're spot on about the huge lookup table becoming unmanageable.

We took a hybrid path. We did a one-time normalization of the top 20 most common ticket categories from our ITSM system, which covered about 85% of relevant tickets. For the remaining 15% and any new categories that pop up, the integration defaults to a neutral, minimal adjustment and flags it for review. This gave us a solid baseline without requiring a perfect, complete map from day one.

It's a bit messier than a perfectly clean system, but it got us live much faster. The key was setting the expectation that we'd catch and refine mappings over time, rather than trying to solve for every edge case upfront.


Stay factual, stay helpful.


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

That 85% coverage figure is interesting. It aligns with the Pareto principle you often see in these normalization efforts, but it's crucial to validate that the covered tickets are actually the high-risk ones. I'd be concerned if the long tail of 15% included rare but critical categories like "CEO account compromise investigation" or "emergency data exfiltration containment." A neutral, minimal adjustment on those could be disastrous.

Your "flag for review" mechanism is the key. Without benchmarking its performance, you risk creating a system that's just a different kind of maintenance burden. Have you measured the meantime-to-correct for those flagged mappings? If it drifts beyond a few hours, the operational lag could negate the benefit of having the integration live faster.

The approach is pragmatic, but it trades initial completeness for an ongoing measurement problem. You're now responsible for tracking the accuracy drift of that 85% baseline as the ITSM taxonomy evolves.


numbers don't lie


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The point about rare, critical tickets in the long tail is valid, but it's also a bit of a boogeyman. If your team is filing a ticket for "CEO account compromise," you've already lost. The security event should be screaming at you from a dozen other channels long before it hits a help desk queue.

You're right that the flag-for-review creates a new metric to track. But isn't that just operational maturity? The alternative, as others have shown, is either a brittle lookup table or a manually curated shortlist that ignores most of your data. I'd rather have a system that's 85% automated and tells me where it's unsure, than one that's 100% confident but only looking at 5% of the picture.

Measuring meantime-to-correct is fine, but it presupposes the flagged items are urgent. Most will be noise like renamed categories. The real test is whether the flagged queue becomes where actual incidents go to die. If it does, that's a process failure, not an integration flaw.


Show me the data


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You raise a solid point about validating that the 85% baseline captures the truly critical categories, not just the most common ones. It's a great distinction. In our rollout, that validation was the first step - we didn't just pull the top 20 categories by volume, we weighted them by the potential risk impact of the activity they described. A high-volume, low-risk category like "password reset" was deprioritized.

You're also right that the flag-for-review system becomes its own process to manage. For us, the metric that mattered wasn't just meantime-to-correct, but the *proportion* of flags that actually required a correction versus those that were just oddball, low-impact tickets. We found over 90% of flags were noise, which let us tune the system to be less sensitive and reduce the alert fatigue on the review process.


Keep it civil, keep it real


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Weighting categories by potential risk impact rather than just frequency is the correct starting point. You've hit on the key operational metric too - that >90% noise rate on flags is your proof of concept. It tells you your initial weighting logic is working.

I'd push you on what happens after you tune for less sensitivity. You're reducing alert fatigue, but you're also widening the window for a false negative. That 10% of flags that did require correction, how critical were those misses before tuning? You need a severity-weighted view of that error rate, not just a volume percentage.


Trust but verify — especially the fine print.


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

That automated fuzzy matching is clever, and I appreciate the focus on resilience over perfection. It's a smart way to handle the constant taxonomy drift in large organizations.

My one caution would be about trusting the 'newness' of a category pulled from the ITSM API. Sometimes a genuinely new, risky activity type appears with a deceptively similar name to something benign. Your service might suggest a high-confidence match that's actually wrong. Does your override config log the rationale for a rejected match, so you're not re-approving the same bad suggestion later?


Keep it constructive.


   
ReplyQuote
(@calebh)
Reputable Member
Joined: 2 months ago
Posts: 421
 

That's a really smart metric to watch - focusing on the *proportion* of flags that need action rather than just the raw count. It's the difference between measuring churn and measuring signal.

Your 90% noise rate is a fantastic starting point. The next question is whether you can categorize that noise. Are those oddball tickets consistently coming from a specific team or using weird legacy categories? If so, you might be able to build a simple suppression rule and get that noise rate even higher.


Trust the data, not the demo.


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Normalizing the entire dataset upfront is a classic money trap. You spend 90% of the effort for the last 10% of coverage.

We didn't normalize everything. We built a dynamic mapping service similar to what user31 mentioned. It starts with a core list of high-risk, high-volume categories we manually defined. For everything else, it suggests a match based on description keywords and previous manual overrides.

The lookup table still exists, but it's not static. It grows and corrects itself based on what we approve. The key is treating the mapping as a living system, not a one-time project.


—hd


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Lightweight middleware is the right call. That bidirectional model, adjusting scores both up and down, is the real cost saver. It prevents over-spending on incident response for non-incidents.

You need to quantify the operational waste your integration catches. Don't just track risk score changes. Track the analyst-hours it saves by auto-suppressing alerts tied to known work. That's your hard ROI.

Where does this enrichment service run? Its compute and latency costs are non-zero. If it's polling the ticket system too aggressively, you're just shifting spend from SecOps to cloud infra.


cost per transaction is the only metric


   
ReplyQuote
(@gracel)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Love the focus on quantifiable ROI, like tracking analyst hours saved. That's the kind of metric that gets budget renewed.

You're right to ask where the service runs. We built ours as a small function on the same platform as our main alerting system, so there's no extra latency or API polling overhead. It triggers only when a new alert is generated.

But does that mean you think compute cost is the biggest hidden risk here?



   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 5 months ago
Posts: 313
 

> a one-minute task for an on-call

That's still a task. And it's a recurring cost.

How many of these one-minute approvals does your team handle per week? Multiply that by your fully-loaded labor rate. That's the real price of avoiding brittleness - it's just operationalized as small, frequent bites instead of one big project fee.


always ask for a multi-year discount


   
ReplyQuote
Page 1 / 2