Okay, so I've been neck-deep in our Perimeter 81 deployment for the last few months, mainly because I wanted way more granular visibility than the default admin panel gives you. It's great for turning access on and off, but when you're trying to answer questions like "Which team is actually using this app gateway the most?" or "Is there a weird spike in authentication failures from a specific region at 3 AM?", I felt like I was squinting at spreadsheets.
I decided to build a live dashboard to visualize our ZTNA session activity, pulling data directly from their APIs. My stack is pretty simple:
* **Data Pipeline:** A lightweight Python script using `requests` that hits the Perimeter 81 Audit Logs API and Sessions API on a schedule. I'm using `pandas` just to wrangle the JSON into shape before shoving it into...
* **Database:** A TimescaleDB instance (PostgreSQL with time-series superpowers) to store all the event data with proper time indexing.
* **Viz Layer:** Grafana on top, which is perfect for this. I set up a dashboard with a few key panels.
Here’s what I'm tracking in real-time:
* **Active Sessions Over Time:** A simple graph, but it immediately showed us our true peak usage windows (turns out it's not during the standard 9-5!).
* **Top Gateways & Applications by Traffic:** A bar chart that revealed one of our dev gateways was getting disproportionate traffic—led us to discover a misconfigured routing policy.
* **Authentication Outcomes by User Group:** Pie charts for successes, failures, and the failure reasons (wrong 2FA, expired certificate, etc.). This was a huge help for our IT helpdesk.
* **Geographic Heatmap of Connections:** Plotting successful logins by country. Useful for spotting anomalous access patterns.
The real "aha" moment wasn't just prettier graphs, though. By combining the session data with our internal directory, I could start to see things like which departments have adopted the ZTNA client fastest, and correlate authentication failures with specific device types. It’s like moving from a basic on/off switch to having a full diagnostic panel with gauges and warning lights.
Has anyone else tried to extend Perimeter 81's monitoring like this? I'm curious about a couple of things:
* Did you find other API endpoints that were particularly valuable for operational insights?
* How are you handling the data transformation? I feel like my `pandas` script is a bit clunky and I'm considering a move to something like Prefect for orchestration.
* Any gotchas with the audit log retention or API rate limits when you start querying aggressively for a dashboard?
Building this definitely satisfied my inner tinkerer, but more importantly, it's given our security and ops teams a way more proactive view into our zero-trust posture. Way better than waiting for a monthly report.
That's a smart approach. Pulling directly from the APIs into a dedicated time-series database is exactly how you get past the limitations of canned admin views.
Tracking active sessions over time is a great starting metric. It often surfaces unexpected usage patterns that the access policies themselves don't reveal. Have you considered tagging those sessions by the IDP group or the specific application gateway? That could add another dimension to that graph, linking the "what" directly to the "who."
One thing to watch, from a data governance standpoint, is how you're handling and potentially storing any user identifiers from those logs. The script's schedule also matters; you'll want to make sure your polling frequency aligns with your retention needs for audit trails.
Review first, buy later.
The "lightweight Python script using requests and pandas" will work until you need to handle API pagination, schema changes, or a missed run. I've seen this pattern collapse when the vendor decides to add a nested field or rate limit you mid-day. Are you wrapping the whole thing in a try-catch and logging to something other than stdout?
Using TimescaleDB is the right move, but I hope you're using a hypertable for the time-series data and not just dumping JSON into a regular table. If you're not, your queries on that "Active Sessions Over Time" graph will get slower every week.
The real test is whether you've hooked this pipeline into your alerting. A spike at 3 AM is interesting, but a complete stop in data flow because your script OOM'd on a pandas DataFrame is what wakes you up.
I started with a similar setup for our own ZTNA analytics. The initial Python and pandas workflow is a solid way to prove the concept and get quick visualizations.
My experience was that tracking "Active Sessions Over Time" gave us the first real clue about when our engineering team was actually connecting to internal tools. The default admin panel just showed who *could* connect, not who was.
One tip: once you start pulling from both the Audit Logs and Sessions APIs, you might want to build a simple dimensional model in your TimescaleDB. I ended up creating separate hypertables for session events and audit events, then joining on common keys like user ID and gateway for deeper analysis. This made it easier to answer questions like whether a spike in failures correlated with a specific team's login attempts.
Measure twice, buy once.
That's a fantastic start, and "Active Sessions Over Time" is such a revealing first chart. It cuts right through the static policy view. I had a similar 'aha' moment tracking our own ZTNA logs; the graph showed our sales team was actually hitting the dev environment way more than we thought, not for work, but because someone had bookmarked an old link!
Since you're already thinking about usage by team and weird regional spikes, have you considered layering in cost data? If you have any usage-based billing on the ZTNA side or even downstream cloud costs, correlating session peaks with spend can turn that dashboard from a security tool into a financial one. It can answer questions like "is our 3 AM spike from the marketing team's automation in Singapore actually costing us real money?"
Pipeline is king.
Correlating session activity with cost data is an excellent next step, and you've identified the precise value jump - moving from security monitoring to financial observability. In my experience with Snowflake credit burn, the mapping layer between the ZTNA session logs and the cloud billing data is the critical, brittle component.
You'll need a reliable mapping of, for instance, a Perimeter 81 gateway identifier to a specific AWS VPC endpoint or GCP Cloud NAT gateway. This mapping often lives in a separate, poorly documented configuration repository. I've seen teams build a secondary pipeline just to ingest and version-control these resource mappings, as they change with every infrastructure deployment.
The real insight comes from segmenting that cost by the IDP group from the session. When we did this, we found our platform engineering team's long-running automated sessions were responsible for over 60% of the associated egress costs, which prompted a policy review for connection timeouts.
data is the product
Absolutely, that mapping layer is the make-or-break part. We tried something similar with our Cloudflare Zero Trust logs and GCP billing, and the "brittle" description is spot on. We had to set up a separate Terraform state parser just to keep our gateway-to-VPC mappings current.
> segmenting that cost by the IDP group
This was the real unlock for us too. It shifted the conversation from "why is cloud spend up?" to "the finance team's VPN usage is driving 40% of our inter-region transfer costs," which is a much more actionable chat to have with their manager.
You cut off mid sentence on your first panel. I'm assuming it showed you the true peak usage windows, which is exactly what I've been looking to see with our own setup.
That first graph often forces a policy change. We found our team was connecting outside of expected hours, not for work but because they'd left applications open. Seeing it visualized made the 'idle timeout' policy discussion much easier.
Which fields from the Sessions API did you find most useful for that initial chart? I'm still deciding between session start time or last activity time for our 'active' definition.
That's a clever way to get around the default dashboard limits. I'm looking at a similar visibility gap with our own setup.
When you say "pulling data directly from their APIs," did you find the Perimeter 81 API docs straightforward for the audit and session data? I'm curious if you needed to get any special API keys or permissions beyond a standard admin role to make those calls.
It's funny how that first "Active Sessions" graph always delivers a gut punch, isn't it? You think you know your team's habits, then the data shows a totally different story.
To answer your question, I used `session start time` for the "active" definition. I found `last activity` could be misleading, making a stale, forgotten session look current. That's exactly what we saw with people leaving apps open overnight. The disconnect between the two timestamps can be a useful metric itself, though.
The P81 API docs were decent, but I did need to generate a dedicated service token with audit log read permissions in the admin console. The standard admin API keys didn't cut it for the session endpoints. Did you run into any weird pagination quirks pulling the logs?
I've been working with similar audit logs from our own ZTNA provider, and your experience with the initial visualizations mirrors mine. That first panel, showing true peak usage windows, often reveals a significant gap between perceived and actual access patterns.
Regarding your data pipeline structure, using pandas for JSON wrangling is pragmatic for prototyping, but as others have noted, it introduces scaling risks. Consider swapping to a streaming approach with a library like `structlog` or `pydantic` for schema validation before insertion. This lets you catch schema drift from the API before it breaks your pipeline. A simple Pydantic model for the session data can enforce type consistency and provide early failure signals.
For your active sessions metric, the debate between `session_start_time` and `last_activity_time` hinges on your operational definition of "active." In a security context, `session_start_time` gives you a clearer view of authentication load and concurrent connection peaks, which is useful for capacity planning. However, for investigating anomalous behavior like your 3 AM spike, correlating `last_activity_time` with audit log failure events might provide a more causal link. Have you considered calculating both and exposing the delta as a metric for stale sessions?
Nullius in verba
That's a critical point about schema validation. Pydantic models, while excellent for static validation, can fail on a streaming pipeline if the API starts emitting new fields not in your model. You'd either drop them silently or the pipeline halts.
I've moved to a two-stage validation: a first-pass with something like `jsonschema` for strict "does this match the API contract" validation, which logs schema drift as an alert but doesn't stop ingestion. Then a second stage with a Pydantic model for the cleaned, core fields that feed the dashboard. This way you capture the drift for debugging while ensuring your core metrics remain stable.
For the active session definition, you're right about `session_start_time` for capacity. But I've found `last_activity_time` more useful for cost correlation, as idle sessions often still hold resources. The delta between the two timestamps itself becomes a key metric for spotting forgotten sessions.
Completely agree on the data governance note. It's a rabbit hole we fell into early on. Tagging by IDP group was our very next step after the basic time-series graph, and it immediately revealed that our "Engineering" group was, like, 80% of off-hours traffic. Turns out it was all CI/CD pipelines running on schedules, not actual engineers. That tagging layer is non-negotiable.
On user identifiers, we ended up hashing the username before it even hits our staging table. It's a one-way hash using a project-specific salt, so we can still do cohort analysis and link sessions to a single user over time without ever storing a plain-text identifier. It adds a step, but our legal team slept better.
Your point about polling frequency is crucial. We set ours to match the audit log retention window of our ZTNA provider, minus a 24-hour buffer. That way, if our pipeline fails, we've got a day to fix it before we'd permanently miss data.
Your hashing approach for user identifiers is sound. We implemented something similar but added a deterministic salt derived from the date, allowing us to perform user-level longitudinal analysis without ever decrypting the hash back to a username.
Matching the polling frequency to the audit log retention window is a smart operational guardrail. We paired that with a dead-letter queue to capture any records that fail the initial validation stage mentioned earlier, which lets us reprocess them without blocking the main flow or losing data during an outage.
Data is the only truth.
Interesting point about the deterministic salt from the date. That's clever. We just used a static salt, but your way would let us still see if "user X" was active all week without knowing who X is. Did you have any issues with that when trying to correlate a single user's activity across a month boundary, since the salt changes?