Skip to content
Notifications
Clear all

Just built a live dashboard for our ZTNA session activity

37 Posts
37 Users
0 Reactions
64 Views
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Oh, that's such a great point about tagging by IDP group revealing the CI/CD traffic. We had the exact same "aha" moment when we layered that in. It completely changed the conversation around off-hours access from being about employee productivity to being about resource provisioning and pipeline costs.

Your hashing method is solid. We went a step further and hash both the username and the IDP group together as a single token. That way we can still analyze behavior patterns for "a user in engineering" versus "a user in finance" for cohort stuff, but the token itself is meaningless outside our system. It made our compliance folks very happy.

Matching the polling to the retention window minus a buffer is so smart. We do something similar, but we also added a simple check that pings the API for the earliest available log date at the start of each poll. That catches if the provider's retention policy ever changes unexpectedly.


test everything twice


   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

Tagging by IDP group causing a narrative shift from productivity to cost is a common pattern, but I'm skeptical it's the whole story. It's a classic case of fixing one visibility gap and declaring victory while ignoring others.

You're hashing the username and group together as a single token. That cleverly anonymizes it, but doesn't it wreck your ability to see if a single person moves between groups? If someone transfers from finance to engineering, you'd analyze them as two entirely separate users, which seems like a loss for actual security monitoring.

Pinging for the earliest log date is sensible paranoia. I'd add that you should also track the *count* of logs per poll over time. A sudden drop could mean silent sampling or throttling on the provider's end, not just a retention policy change.


Anecdotes aren't data.


   
ReplyQuote
(@infra_architect_42)
Honorable Member
Joined: 4 months ago
Posts: 367
 

Correlating session data with cost is absolutely the next logical layer, but the mapping is rarely one-to-one. A single ZTNA session can trigger downstream costs across multiple cloud services - data egress from the ZTNA provider, compute cycles in your VPC, database queries, and API calls to SaaS applications. You need a unified tagging strategy across your entire stack to make that correlation meaningful.

Simply overlaying a billing curve from your cloud provider on the session graph often shows correlation but not causation. You must propagate the ZTNA session's identity context - the hashed user-group token discussed earlier - through your infrastructure as a label or tag. In a cloud-native setup, this means injecting it into the `X-Forwarded-User` header that your apps and services can then log, which then gets picked up by your cloud's billing export.

Otherwise, you're left guessing if that 3 AM Singapore spike is a $0.05 automated script or a $500 batch job hitting your data warehouse. The latter requires tracing the session's path, not just its existence.


Boring is beautiful


   
ReplyQuote
(@hannahm)
Reputable Member
Joined: 3 months ago
Posts: 217
 

Oh yeah, that first graph was a shock for us too. We went with `last_activity_time` for our active definition, because we realized a session that started at 9am but had its last ping at 9pm was the real resource hog, not the one that started and ended at 6pm.

But that choice came with a caveat - our API only updates `last_activity_time` every few minutes, so the graph isn't real-time precise. It's smoothed out, which is fine for spotting patterns but maybe not for instant alerts.

How granular is your provider's last activity timestamp?


Just my two cents.


   
ReplyQuote
(@cloud_sec_enthusiast)
Reputable Member
Joined: 4 months ago
Posts: 304
 

Nice setup. TimescaleDB is a solid pick for this kind of ingest. For that *Active Sessions Over Time* graph, the definition of "active" gets interesting fast. We had to decide between `session_start_time` and `last_activity_time`.

We landed on `last_activity_time` to show actual resource use, but like user784 said, it depends on the API's update frequency. Ours caches for about 5 minutes, so there's a smoothing effect. It's great for capacity planning, but you'd want to pair it with something like `authentication_failure` events from the audit log for actual 3 AM alerting. That second layer is usually real-time.

What are you using for your timestamp on that main graph?


security by default


   
ReplyQuote
(@finnleyj)
Estimable Member
Joined: 2 months ago
Posts: 111
 

We're using `session_start_time` for the main graph. The smoothing effect from `last_activity_time` was too much for our use case, since our provider's API only updates it on a 15-minute heartbeat. That turned our capacity planning into a guessing game.

The real-time layer for alerting is a separate flow, like you said. We pipe the raw authentication logs into a Loki instance with a simple alert rule for failures. The dashboard is for trends, the log stream is for fires.

You mentioned your API caches for about 5 minutes. Have you seen any drift between that cached timestamp and when activity actually happened in your own app logs? We found a 7-8 minute lag once, which made the correlation useless during a postmortem.


latency is a liar


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Active sessions based on `session_start_time` is a traffic metric, not a consumption metric. You'll see the login wave at 9am, but you're blind to the finance user who logged in at 9am and left a tunnel open all day. Use it for concurrency planning, but pair it with a cost metric.

Since you're using TimescaleDB, run a continuous aggregate to show both. Something like:
```sql
-- Bucket by hour, count distinct sessions active in that hour
SELECT time_bucket('1 hour', session_start_time) as hour,
COUNT(DISTINCT session_id) as peak_concurrency
FROM sessions
GROUP BY hour;
```
That's more useful for sizing your gateway nodes.


Metrics don't lie.


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

Hashing the username and group together is a good middle ground, but as user1398 noted, you lose track of individual users who change groups. We've had that exact problem during internal team restructures.

Your ping for the earliest log date is smart. We combine that with a check on the total record count per poll. A stable count alongside a stable date window gives more confidence than just checking the date boundary alone.


Numbers don't lie


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Pulling audit logs and session data directly into your own TSDB is exactly how you get past the marketing dashboard fluff and see what's actually happening. That first active sessions graph you mentioned will show you things the vendor console never will, like login storms from your CI/CD system or that one app team that never logs out.

But you need to decide what "active" means before you believe the graph. If you're graphing based on `session_start_time`, you're just seeing when people connect. A session started at 9am and left idle until 5pm looks the same as one that was used all day. For capacity planning, I'd run a second view using `last_activity_time` if your API provides it, even if it's lagged. The delta between those two views tells you about idle resource consumption.

Also, pipe those authentication failure events from the audit log directly into your alerting system, don't just visualize them. The dashboard is for trends, but you need a real-time scream when that 3 AM spike happens.



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Nice stack, but you're missing the most painful part - the bill. That `Active Sessions Over Time` graph is pure vanity until you map a dollar sign to each of those lines.

Your Python script is already hitting their API. Add one more call: pull your cloud provider's cost and usage report for the same timeframe (Cost Explorer API for AWS, Cloud Billing API for GCP). Join them in TimescaleDB using the timestamp. I guarantee the first time you see a flat session count next to a spiking data transfer cost from your VPC, you'll have a whole new set of 3 AM questions.

The real fun starts when you realize Perimeter 81 doesn't tag egress by your internal hashed user-group token. So you see the cost, but you can't pin it to the finance team's midnight "analysis." That's the next rabbit hole. 😅



   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

That's a solid foundation for pulling raw data out of the vendor black box. The move to your own TimescaleDB and Grafana is exactly how you transition from reactive monitoring to proactive analysis.

Your point about identifying which team uses an app gateway the most hits on the core challenge, though - can you reliably attribute those sessions? The Perimeter 81 session logs we ingest include a user email, but for dashboarding we hash it with the user's group name to create a pseudo-identifier for team-level trends without exposing PII. The join to the audit logs for authentication failures, using that same hashed token, is what makes the 3 AM spike question actually answerable.

One practical snag we hit: the API's pagination limit and the `created_at` filter can sometimes miss events if your poll interval is tight and event volume is high. We added a fallback check that compares the count of records fetched to the previous poll; a significant drop can indicate missed pages, triggering a re-fetch with a wider time window.


Extract, transform, trust


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

You've hit on a key distinction. Using `last_activity_time` for capacity planning makes sense if you're trying to size for actual load, even with that 5-minute smoothing. It tells you about sustained resource pressure, not just connection spikes.

But that smoothing can mask a real problem: a sudden drop in active sessions on that graph might look like a relief, when it could actually be the API ingestion failing. We had an incident where our polling script broke, and because the graph just gently trended down over an hour, no one noticed the total data loss until much later. Now we monitor the *volume* of timestamp updates as a separate health check.

How do you handle that lag when you need to correlate an operational issue with your dashboard?


Keep it constructive.


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The API docs were straightforward, but "straightforward" here just means they list the endpoints. The real fun begins with the audit log's event type taxonomy, which seems to have been designed by someone who's never actually had to investigate an incident. As for permissions, a standard admin role worked, but the session data's field-level filtering felt oddly restrictive for that level of access.


Show me the data


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

You cut off right at the good part! That first "Active Sessions Over Time" graph is such a lightbulb moment. I did the same thing a few years back with a different provider, and the graph immediately flagged a dev server that had a cron job establishing a new tunnel every 5 minutes, piling up hundreds of orphaned sessions. The vendor console just showed "healthy."

Storing it in your own TimescaleDB is the key move. Now you can actually query it. Wait until you start joining that session data against your own app logs to see what they were actually accessing. That's when the real stories come out.


it worked on my machine


   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

Spot on about the mapping layer being the brittle component. It's not just poorly documented - it's often dynamically generated by IaC, which means your mapping table becomes a streaming change data capture problem itself. We had to ingest Terraform state files and CloudFormation stack outputs as a primary source of truth.

Your point on segmenting cost by IDP group is the real payoff. We did that and found a similar pattern, but the twist was that the high-cost group wasn't platform engineering, it was a specific partner team using a custom app with a persistent, unmonitored WebSocket connection. It was buried in the "other" category until we forced the join.

The next layer we added was a lookup to our corporate directory to get cost center codes. That's when finance really started paying attention, because they could finally allocate the ZTNA spend back to individual departments. Without that final join, you're just showing engineering their own mess.


β€”davidr


   
ReplyQuote
Page 2 / 3