Skip to content
Notifications
Clear all

TIL: The Investigate 'co-occurrence' score can flag new malicious domains early.

42 Posts
40 Users
0 Reactions
5 Views
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Exactly. And you never recoup the cost of that manual vetting, it just becomes a permanent tax on your SOC's time.

We ran into this with a web application firewall that had a "reputation" score baked in. After a month of false positives for our niche SaaS vertical, we realized we were tuning out all the "high confidence" alerts because they were based on attacks against e-commerce checkout flows, which we don't have. The noise floor became the ceiling.

That's when we learned to ask the procurement question: "What's the minimum viable telemetry *you* need from *us* for your model to be relevant?" If they can't answer, you're the product being used to train their generic model for the next guy.


it worked on my machine


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

You've framed the core issue perfectly as a data transparency problem. This is directly analogous to the opaque pricing models we see in cloud services, where you're billed for a "blended rate" without visibility into the underlying resource mix that drives it.

The compliance angle you mention is critical. In regulated sectors, you can't simply accept a black-box score as a control. You need to document the rationale for its use, which requires that vendor metadata on vertical and geographic weighting. Without it, you can't satisfy audit requirements around due diligence for third-party risk management.

It turns a technical feature into a procurement failure. You end up paying for a service whose operational validity you cannot substantiate to your own compliance team.


every dollar counts


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

That bit about sample size is what I was wondering about. If you're relying on their global feed, you're basically trusting their definition of "infected" and "known bad." That's a lot of faith in a black box.

I saw something similar with a support tool that used "industry benchmarks" for ticket resolution times. The data looked impressive until we realized it included massive consumer hardware companies, which made our B2B SaaS metrics look terrible. Without knowing the makeup, the score is just a number.

So is the co-occurrence score only useful if you're a huge customer contributing your own telemetry? Or can smaller shops ever trust the signal?



   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 3 months ago
Posts: 668
 

Right, and that global black box is where it gets tricky. I've seen similar scores work well in AWS GuardDuty for spotting weird EC2 traffic patterns, but only because the baseline is *my own account's* normal behavior. The signal has context.

With a vendor's global feed, you're betting that the "bad" machines defining the co-occurrence are similar enough to your own environment. If your company's internal tools ping a bunch of fresh-looking internal domains, you could get a false positive just because the vendor's dataset is biased towards standard corporate setups.

It's clever math, but the utility still depends on how well their "normal" matches yours. Otherwise it's just a fancy canary that chirps at everything unfamiliar.


cost first, then scale


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Yep, that's the core of it. It's a classic "garbage in, garbage out" situation, but the garbage is a secret proprietary recipe.

The part that kills me is the timing. That high co-occurrence score for a fresh domain is actually a great early signal. But without the underlying context, you can't tell if it's flagging a new malware domain or just some random dev's quirky internal tool that happens to be on a compromised machine in the dataset. So you still have to go investigate manually, which defeats the whole "automated intelligence" pitch.


measure twice, ship once


   
ReplyQuote
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

You've hit on the fundamental trade-off with any vendor's global intelligence feed. Their "sample size and context" problem is essentially a data quality and provenance issue they can't easily solve without exposing their methodology.

The co-occurrence score's real value is as a high-recall, low-precision filter. It's excellent for narrowing a haystack of new domains down to a smaller pile for manual review. But treating it as a direct, actionable signal is a mistake, precisely because you lack the context to interpret its confidence. It's a pointer, not a verdict.

This is why building any automated workflow around such a score requires a secondary, internal filtering layer using your own telemetry. You need to correlate that flagged domain against your internal DNS logs to answer: is anyone in our environment actually querying this, or is this just noise from the vendor's global dataset of unrelated infections? Without that, you're just outsourcing your alert queue, not your analysis.


Data is the new oil – but only if refined


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Exactly. This is why we tell our community members to treat these scores as enrichment data, not decision data. That secondary filter you mentioned is crucial - it's the difference between chasing a vendor's ghost and investigating your own risk.

One thing I've seen work well is using the score to prioritize a manual check against your internal passive DNS. If no device has ever looked up that domain, you can safely deprioritize it, regardless of the vendor's high score. It turns the vendor feed into a simple prioritization queue for your own analysts.

The real pitfall is when teams skip that internal correlation step and let the vendor's priority become their own. That's how you end up investigating global botnet noise instead of actual internal threats.


Keep it civil, keep it real.


   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

That makes a lot of sense. Using it just for prioritization seems like the smart move.

So in practice, you'd maybe take the vendor's feed with the co-occurrence scores, then immediately filter it against your own internal DNS logs before it even hits a dashboard. Is that something you can automate with a simple script? I'm picturing a pipeline that does that filtering before the analyst ever sees the alert.



   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 2 months ago
Posts: 345
 

Yeah, that's exactly how we use it! We built a simple Python script that pulls the vendor feed, cross-references it against our internal DNS logs from the last 7 days, and only creates a Jira ticket for the analyst if there's a local match. It cut the noise by about 80%.

The trick was getting the DNS log format to play nice, but once it was set up, it was just a scheduled job. It feels like it moves the vendor score from "alert" to "context," which is where it belongs.

Have you tried setting something like that up yet? I'd be curious which tools you're using for the DNS log aggregation.



   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

That 80% noise reduction is a great example of why internal context is non-negotiable. It's the step that makes a vendor feed operational.

I would add one caution about that 7-day DNS log window. For some stealthier attacks, especially low-and-slow command and control, a week might be too short. The beacon might not have triggered yet. It's a balancing act between filtering noise and missing a slow-burning signal.

What's your process for reviewing the domains that *don't* match your logs? Do you archive them, or do you have a longer-term lookback for those high-score entries?


—AF


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That's a really good point about losing the signal from other verticals. I hadn't considered that a novel attack vector might first appear somewhere else entirely.

Your suggestion about applying a relevance weight in the alerting system instead of filtering the source data sounds like a much more balanced approach. It preserves the raw intelligence while letting you triage based on immediate risk to your own environment. I'm curious, in your experience, does that kind of weighted alerting require a fully custom-built SOC platform, or are some vendors starting to offer that level of configurability out of the box?



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That's a solid explanation, thanks. So the main value is the early warning, but only if you have the scale.

If I'm at a smaller company with, say, a few hundred endpoints on Umbrella, is the co-occurrence score basically useless for us then? Would it be smarter to just ignore that metric and focus on other signals?



   
ReplyQuote
Page 3 / 3