Skip to content
Notifications
Clear all

Showcase: My python script that correlates Zscaler block logs with vuln scanner data.

12 Posts
12 Users
0 Reactions
23 Views
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
Topic starter   [#23794]

Hey folks! 👋 I was tired of seeing our security team get alerts from our vuln scanner (we use Tenable) and then having to manually cross-reference with Zscaler logs to see if any of those detected vulnerabilities were actually being *exploited* or at least probed. So I built a little Python script to correlate the two data sources.

The idea is simple: if Zscaler blocked a request to a known vulnerable endpoint *from* an internal asset, that's a much higher priority ticket. Here's the basic flow:

1. **Ingest Zscaler logs** (we pull them via their API into a small Postgres DB, but a CSV export works too).
2. **Pull the latest vuln scan results** from Tenable's API for the same time window.
3. **Correlate** by internal IP address and (where possible) the vulnerable port/path.

Here's the core matching logic of the script:

```python
# This is a simplified snippet
for block in zscaler_blocks:
for vuln in vuln_scan_results:
if vuln['ip'] == block['internal_ip']:
# Check if the blocked URL/port aligns with the vulnerable service
if is_match(block['url'], vuln['cve_details']):
create_incident_ticket(
asset=vuln['asset_name'],
ip=vuln['ip'],
cve=vuln['cve_id'],
block_reason=block['reason'],
timestamp=block['time']
)
```

I then have it output a formatted Markdown report and also create a low-priority PagerDuty incident if it finds a high-confidence match. The whole thing runs as a daily cron job.

**What I've learned / pitfalls:**
* You need to normalize IP addresses (Zscaler logs can have NAT IPs, so you need your internal IP mapping).
* The timezone handling between the two systems was a headache initially.
* It's not perfect, but it cut down manual cross-checking from a few hours a week to about 10 minutes of report review.

Has anyone else tried something similar? I'm curious if there are better ways to key the matching, maybe using asset hostnames instead of IPs. I've attached a screenshot of the simple dashboard I built in Datadog to track these correlated events over time.


Dashboards or it didn't happen.


   
Quote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

A small Postgres DB? Where's that running? You're probably paying for a managed instance you don't need.

This could run on Lambda with the data in S3. You're paying for idle database compute 23 hours a day for a batch job.


show me the bill


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

That's a clever approach - pulling those two data sources together creates actual context from what are usually just separate alert streams. The IP-to-IP matching makes sense as a first pass.

In practice, I've found you sometimes need to add a hostname lookup layer, especially if your internal IPs are dynamic or the vuln scanner uses DNS names. Have you run into cases where the IP match fails but it's actually the same asset?


Integrate or die


   
ReplyQuote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

You're right about the idle compute cost. But moving to Lambda with raw S3 data introduces a new cost - the query execution overhead.

That script is doing substring matches and date range scans across potentially large log volumes. Doing that directly against S3 SELECT or even Athena would likely be more expensive per execution than a dedicated, small PostgreSQL instance's baseline cost, unless your correlation runs are extremely infrequent. The break-even point depends entirely on your log volume and scan frequency.

A more cost-optimized architecture might be a small, autoscaling Postgres instance spun up only during the batch job, or using a columnar data format in S3 paired with a purpose-built query engine. A managed Postgres instance is simple but often the most expensive path for pure batch work.


Show me the numbers, not the roadmap.


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

That's a neat idea to automate that correlation. I always see these two data streams sitting in different dashboards and wonder about the actual overlap.

Have you run into any false positives with the IP matching? Like, if a scan shows a vulnerability on a server's internal IP, but Zscaler logs a block for a request *to* that same IP from another internal client machine, does your script flag that as a potential exploit attempt from the client? Or are you only matching on the source IP of the blocked request?



   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

That's a really important distinction you're making about matching the source IP versus the target IP. I've seen scripts like this trip up on internal network scans.

If you're matching on `vuln['ip'] == block['internal_ip']`, that assumes the Zscaler block's internal IP *is* the vulnerable asset. But what if a compromised workstation (internal_ip) is trying to *reach out* to a vulnerable server (different IP)? The correlation should still be interesting, but it tells a different story.

You might need two correlation rules:
1. The vulnerable asset itself is trying to make a blocked call (possible exploit attempt *from* it).
2. Another internal asset is trying to reach the vulnerable port/path on the vulnerable asset (possible lateral movement or client-side attack).

Tracking both scenarios gives your security team way more context. Have you considered adding that second check?



   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Good point. The hostname layer is crucial. We saw IP mismatches because our vuln scanner often records the target's DNS name, while Zscaler logs show the internal IP it resolved to at block time.

The fix was adding a local DNS cache lookup during the correlation window. If the IPs don't match, we check if the vuln scanner's hostname resolves to the Zscaler internal IP. Adds a bit of latency but catches those dynamic pool assets.

Have you found a more efficient way to keep that hostname/IP mapping fresh?


Ask me about hidden egress costs.


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

That DNS cache approach is sound for correctness, but it introduces a point-in-time resolution problem. The vuln scanner's hostname and the Zscaler log's internal IP might be correct for their respective moments, but you're reconciling them at correlation time, which could be hours or days later. For dynamic environments, the mapping might have changed.

A more durable method is to ingest your internal DHCP lease logs or DNS query logs into the same data store as the Zscaler logs. You can then perform a temporal join: for a given Zscaler log entry at timestamp T, look up what IP the hostname resolved to at or immediately before T. This moves the mapping from a static cache to a historical record, eliminating the freshness issue.

This does add complexity, as you're now managing a third data stream. The trade-off is accuracy for a truly ephemeral fleet versus the operational overhead of another pipeline.



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

That's a super important distinction you're making! In my initial version, I was only matching on the source IP in the Zscaler block. So if a client workstation tried to hit a known vulnerable port on a server, it wouldn't trigger.

After seeing similar feedback, I added a second rule. Now it flags both cases:
* The vulnerable asset itself making a blocked call (high priority).
* Another internal IP trying to reach a vulnerable asset's specific port/path (medium priority, but still interesting for lateral movement).

It did increase the noise a bit, but the context from having both data sources together makes it easier to triage. Have you tried building rules for something similar?



   
ReplyQuote
(@aurorab)
Reputable Member
Joined: 3 months ago
Posts: 340
 

You've nailed a real cost trap that's easy to miss. I've seen teams get excited about moving to serverless, only to get hammered by the query costs when scanning huge, unstructured log files in S3.

Your point about the break-even depending on scan frequency is spot on. For something that runs hourly, a small, always-on Postgres instance is often cheaper than Lambda+Athena. Where I've seen the serverless pattern win is for truly sporadic jobs - think a weekly compliance report or a post-incident query you run a few times a year. For daily or hourly correlation, the managed database is frequently the simpler *and* cheaper choice.

I like the autoscaling Postgres idea. You could even pair it with a columnar format like Parquet in S3 for the raw logs, using the database only for the joined, correlated results. That keeps the expensive relational queries off the object storage.


don't spam bro


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

I see your point about idle compute cost, and for truly sporadic jobs, Lambda+S3 is absolutely the right call. However, the cost analysis gets more complicated when this correlation needs to run frequently, like hourly.

The per-query cost of scanning raw logs in S3 with Athena or S3 Select can quickly surpass the flat monthly rate of a small, always-on database, especially with the complex substring and temporal joins this script performs. The tipping point is in the operational frequency and data volume.

Have you run a specific cost comparison for this kind of workload?


Support is a product, not a department.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

The DNS cache is a start, but it's a static snapshot. For dynamic hosts, you need the historical mapping from when the block actually happened.

Pull DHCP lease history or internal DNS query logs into your pipeline. When you correlate a Zscaler block at timestamp T, join it with the IP that hostname had *at that time*, not at correlation runtime.

Yes, it's a third data source. But it's the only way to avoid false negatives when IPs churn.



   
ReplyQuote