Skip to content
Notifications
Clear all

Showcase: Grafana alerts from Zscaler logs for anomalous location logins

10 Posts
10 Users
0 Reactions
17 Views
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
Topic starter   [#23646]

We built a detection rule for anomalous location logins using Zscaler NSS logs pushed to Grafana Loki. The goal was to flag user logins from geographic locations they've never been seen in before, without waiting for a SIEM correlation.

Key components:
* Zscaler NSS logs (specifically `ZSCALER_NSS_FIREWALL_LOGS`) streamed to a Loki instance via Fluent Bit.
* A Grafana `sql` alert rule that runs a query to find new, unique country codes per user.

The alert query logic looks for any user+country combo that hasn't been seen in the last 30 days. We tuned it to ignore our common corporate VPN exit nodes.

```sql
WITH current_logs AS (
SELECT DISTINCT
parsed.user_login as user,
parsed.country_code as country
FROM logs
WHERE parsed.event = 'auth'
AND parsed.country_code != ''
AND $__timeFilter(timestamp)
),
historical_logs AS (
SELECT DISTINCT
parsed.user_login as user,
parsed.country_code as country
FROM logs
WHERE parsed.event = 'auth'
AND parsed.country_code != ''
AND timestamp >= NOW() - INTERVAL '30' DAY
AND timestamp < $__timeFrom
)
SELECT
cl.user,
cl.country
FROM current_logs cl
LEFT JOIN historical_logs hl ON cl.user = hl.user AND cl.country = hl.country
WHERE hl.user IS NULL
AND cl.user NOT LIKE 'svc-%'
```

Alert fires, sends a notification to our security channel with the user and country. Team can then verify or dismiss. Works faster than our old SIEM workflow.

Biggest pitfall was log volume—had to adjust the Loki retention and chunk size. Also, Zscaler's internal service IPs sometimes get mapped to unexpected countries; we maintain a small exclusion list for those.

—cp


—cp


   
Quote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

Interesting approach using Loki and Grafana for this. The 30-day historical baseline is a practical choice, but have you considered the risk of seasonal employees or contractors who might legitimately log in from a new country after a longer period of inactivity? Your rule would flag them.

Also, the static exclusion for corporate VPN nodes is a good start, but it requires manual upkeep. A drift in those IP ranges could cause false negatives. You might want to layer in a secondary check, like looking for concurrent logins or impossible travel if you ever get timestamps with enough precision.


—at


   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

Your static VPN exclusions won't scale. I'd move those to a separate lookup table your query can reference, maybe in a small PostgreSQL sidecar. Update once, query references it.

Also, 30-day baseline is too short for some roles. We use a 90-day window and still flag for manual review on first alert. Cuts noise.



   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

Interesting technical approach, but I can't help wondering about the bill for that Loki instance. You're running a DISTINCT query over 30 days of firewall logs for every alert evaluation. That's scanning a lot of data, and Loki's query pricing model isn't exactly built for cheap repetitive historical lookups.

You'd probably save a significant chunk by moving the historical baseline logic to a tiny, purpose-built table. Something that gets updated incrementally when new country combos appear, then your alert query just checks the last hour of logs against that table. You're paying for CPU cycles to re-calculate the same DISTINCT sets every time the alert runs.


pay for what you use, not what you reserve


   
ReplyQuote
(@danielz)
Estimable Member
Joined: 2 months ago
Posts: 171
 

You're right about the scanning cost. Loki's query model bites you hard on repetitive full-period scans.

But moving baseline logic to a side table introduces state you have to manage and keep in sync. Now you've got an ETL job and another potential point of failure. Sometimes the brute-force query, even if inefficient, is simpler to reason about operationally.

The real issue is using Loki for this pattern at all. It's a logs database, not for this kind of historical lookup. If cost is a concern, you picked the wrong tool.


show me the logs


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

Yeah, the operational simplicity argument really hits home for me. I'm currently drowning in a side table system I built for something similar - it's a constant "why is this sync job stuck" fire drill.

But you're also right about Loki being the wrong tool here. It feels like we're all trying to force logs databases to act as stateful alert engines because they're what's already there and hooked up.

So what *is* the right tool for this specific job? Something that can cheaply store the "user+country last seen" state and check new events against it. I'm picturing a small key-value store, but that's just another piece of infra to manage. Is there a middle ground?


null


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

The operational simplicity you've achieved with a single, self-contained query is definitely valuable. It's a clear, auditable detection rule that doesn't depend on a separate pipeline's health. However, building on the cost concerns others have raised, you might consider a hybrid approach that retains simplicity while reducing scan volume.

Could you adjust the `historical_logs` CTE to query a much smaller, pre-aggregated data set? For instance, you could run a daily query that materializes the distinct user+country pairs from the previous day into a dedicated Loki label or a separate, low-cost object store table. Your alert rule would then join against that compact daily snapshot for the last 30 days instead of the raw logs. This would cut the repetitive scanning of the full log volume while keeping the logic largely within Grafana's purview.

It shifts the maintenance burden from managing a live sync job to a scheduled aggregation, which often fails more gracefully. If the daily job breaks, you're only missing updates to the baseline, not breaking the alert rule itself, and you'd likely catch it within 24 hours.


Let's keep it constructive


   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 198
 

That's a smart middle ground. Shifting to a scheduled aggregation does sidestep the live sync fragility.

One nuance: the daily snapshot approach still requires managing a separate storage target and a scheduled query. If that's already in your wheelhouse, it's solid. But if you're in a smaller shop where the original appeal was having *everything* in Loki, you've now added two new operational pieces instead of one.

You might also consider if your Loki setup supports recording rules. You could potentially materialize that daily distinct set as a new metric in Prometheus, then query it from there. Keeps it in the observable stack, but now you're crossing systems.


- GG


   
ReplyQuote
(@consultant_carl)
Honorable Member
Joined: 6 months ago
Posts: 412
 

Exactly, that's the pragmatic trade-off. Shifting to a scheduled aggregation turns a live sync problem into a batch job problem, which is almost always easier to troubleshoot and more forgiving when it hiccups.

My caveat to your point about it failing gracefully is around alert fatigue. If that daily job breaks and you're missing baseline updates for 24 hours, you'll get a flood of false positives for *every* legitimate login because the system thinks every country is new again. You need a separate monitor on the aggregation job itself, or you'll train your team to ignore the alerts.

So you trade pipeline fragility for potential alert storms, but at least the storms are visible and you know the root cause immediately. I'd take that deal.


Implementation is 80% process, 20% tool.


   
ReplyQuote
(@gracej77)
Honorable Member
Joined: 3 months ago
Posts: 444
 

You've hit on the key trade-off perfectly. That separate monitor for the aggregation job is non-negotiable, or you'll indeed nuke your team's trust in the alert.

One extra consideration: while a failed daily job causes a storm of *false positives*, a live sync pipeline breaking can cause *false negatives* where new malicious logins get missed entirely. I'd argue the alert storm is the less dangerous failure mode, because it's loud and obvious. Silent failures are scarier.


Keep it real, keep it kind.


   
ReplyQuote