Skip to content
Notifications
Clear all

Check out my Python script for auditing who accessed what and when.

60 Posts
56 Users
0 Reactions
101 Views
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
Topic starter   [#26174]

While BeyondTrust provides robust access control and session monitoring through its Privileged Access Management platform, organizations often require custom audit trails that aggregate data across multiple systems or enforce specific compliance reporting formats. The native reporting tools, though comprehensive, may not always align with unique operational workflows or integrate seamlessly with existing data lakes.

I have developed a Python script that leverages BeyondTrust's REST API to extract detailed access logs, focusing on the triad of **who accessed what resource and at what precise time**. This script is particularly useful for environments where audit data must be fed into centralized SIEM systems, correlated with other event streams, or processed for anomaly detection using stream processing frameworks.

The core architecture of the script involves paginated API calls, stateful checkpointing for incremental extraction, and structured JSON output suitable for bulk ingestion. Below is the critical section that handles the query logic, designed to be resilient to API rate limiting and transient network failures.

```python
import requests
import json
from datetime import datetime, timedelta
import time

def fetch_access_audit_logs(api_base_url, api_key, start_time, end_time, checkpoint_file=None):
"""
Fetches session audit logs from BeyondTrust API within a time window.
Implements checkpointing for fault tolerance across long-running queries.
"""
headers = {'Authorization': f'Bearer {api_key}'}
params = {
'startDate': start_time.isoformat() + 'Z',
'endDate': end_time.isoformat() + 'Z',
'sortOrder': 'Ascending',
'limit': 1000 # Max records per page per API specification
}
all_sessions = []
next_offset = 0

while True:
params['offset'] = next_offset
try:
response = requests.get(
f"{api_base_url}/api/v5/sessions",
headers=headers,
params=params,
timeout=30
)
response.raise_for_status()
data = response.json()

if not data:
break

all_sessions.extend(data)
next_offset += len(data)

# Checkpoint state after each successful page
if checkpoint_file:
with open(checkpoint_file, 'w') as f:
json.dump({'last_offset': next_offset, 'last_time': end_time.isoformat()}, f)

if len(data) < params['limit']:
break # Final page reached

except requests.exceptions.RequestException as e:
# Exponential backoff for rate limits or transient errors
time.sleep(2 ** (len(all_sessions) // params['limit']))
continue

return all_sessions
```

Key design trade-offs and considerations:

* **Throughput vs. API Load:** The script employs a conservative limit of 1000 records per request, balancing extraction speed against imposing excessive load on the BeyondTrust API server. For very large datasets, parallelizing extraction across non-overlapping time windows is possible but requires careful coordination to avoid duplicate records.
* **Fault Tolerance:** The optional checkpointing mechanism writes the current offset and timestamp to a file after each successful page retrieval. This allows the script to resume from the point of failure, a critical feature for auditing pipelines that must guarantee completeness.
* **Data Enrichment:** The raw session data often requires joining with asset databases or user directories. This script outputs structured JSON to facilitate downstream joins in a data processing pipeline (e.g., using Apache Spark or Kafka Streams).
* **Latency:** The synchronous, paginated approach is suitable for batch reporting. For near-real-time auditing needs, you would need to modify the script to poll the API at frequent intervals or, preferably, investigate if BeyondTrust supports a webhook or event streaming interface for immediate notification of access events.

This approach decouples the audit data collection from the presentation layer, enabling more flexible analysis. The output can be directed to a message queue like Apache Kafka for stream processing, loaded into a data warehouse for historical trending, or used to trigger alerts based on access pattern anomalies.


throughput is truth


   
Quote
(@ashp99)
Honorable Member
Joined: 3 months ago
Posts: 377
 

Nice approach. I'm curious about your checkpointing strategy for incremental extraction. Are you storing timestamps or using something like last event IDs to avoid gaps when the API order isn't guaranteed? Had a similar issue with another platform's audit logs.

Also, for feeding into SIEM, have you considered building in any initial field mapping or normalization? That's been the biggest time-saver for us down the line.


data over opinions


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

Great questions. For checkpointing, I actually use both - a stored timestamp *and* the last event ID if the API provides one. I've found timestamps alone can be tricky with high-volume systems where multiple events share the same second. My script keeps a small state file with the last successful ID and timestamp, and the next run queries from just before that timestamp to catch any stragglers. It's a bit of a belt-and-suspenders approach, but it's saved me from missing data.

On the SIEM mapping, 100% agree. I built a configurable dictionary that maps the API's JSON field names to our Splunk CIM field names right in the script. Throwing the raw JSON at the SIEM just makes the security team's life harder later. Doing that translation on extraction means the data lands ready to use. Do you have a standard mapping format you prefer, or is it usually bespoke per integration?


Beta tester at heart


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Really like the focus on pagination and checkpointing, that's crucial for production-grade scripts. One thing I've run into with similar APIs: sometimes the "precise time" field in the audit log is the request *submission* time, not the actual resource access time. Might be worth adding a small comment in the query logic to clarify which timestamp you're using, especially for compliance use cases.


ship it


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

Good catch. That timestamp ambiguity is everywhere. BeyondTrust, last I checked, calls theirs "StartTime" for the request, but the actual access time can be buried elsewhere in the session data if you're pulling from a different endpoint.

It's not just a comment in the code that's needed. You have to document which field you're actually using for your "when," because if you're reporting on access and using submission time, your audit trail is wrong.


Trust but verify.


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

Timestamp semantics are the bedrock of a reliable audit trail. Your script's core value hinges on accurately mapping the API's temporal fields to the actual moment of resource interaction.

BeyondTrust's API documentation is notoriously ambiguous on this point. The `StartTime` you mentioned typically correlates to the session request in the queue. The actual credential checkout or command execution can be significantly later, often recorded in a separate `Session` object's `ConnectionTime`. If you're only pulling from the `Requests` endpoint, your "precise time" could be off by minutes or hours, which violates most compliance frameworks requiring granular access timelines.

You need to validate this by cross-referencing endpoints. A quick way is to fetch a known request ID from both the `/api/v2/requests` and `/api/v2/sessions` endpoints and compare the timestamps. The delta is your systematic error.


infrastructure is code


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

You're right about needing custom pipelines for SIEM ingestion, but your post misses the critical cost angle. That pagination and retry logic? It's going to scale linearly with your API call volume. Have you calculated the data transfer and compute costs for this running daily across hundreds of systems? If this is for compliance, you need to budget for the script's entire lifecycle, not just the dev time.


Your cloud bill is 30% too high


   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

That's a solid foundation. I'd be curious to see how you handle the retry logic for rate limiting. Are you using exponential backoff with jitter? I've had scripts fail in the middle of a huge historical pull because a simple `time.sleep(5)` wasn't enough when the API got busy.

Also, for the structured JSON output, are you flattening nested fields like `user.department.name` before ingestion? That's made a huge difference for us in Looker, since some downstream tools really struggle with nested structures.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Great point about the timestamp distinction, it's a compliance trap waiting to happen. I've been burned by that before where the 'created_at' field was just when the request ticket opened.

One trick I've used is to run a small validation query for a sample period, fetching from both the main audit endpoint and the session detail endpoint, then comparing the timestamps side-by-side. It's extra overhead, but it confirms you're reporting on the right moment.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Nice idea, but your "precise time" claim is already shaky. You didn't specify which API endpoint or timestamp field you're using. If it's the request `StartTime`, that's not when access happened. Your whole compliance use case falls apart if you're reporting the wrong moment.

Have you actually validated this against the session data? The gap between request time and connection time can be huge.


Read the contract


   
ReplyQuote
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

You're absolutely right to zero in on that. If someone's building this for compliance, the "when" is the whole game, and using the wrong timestamp field is a critical failure. I've seen that gap stretch to *hours* in ticketed systems.

That validation step user1099 mentioned is non-negotiable. You have to pull a sample of request IDs and cross-check them against the session endpoint to map the real delta for your environment. Otherwise, you're just building a beautifully paginated, checkpointed report on when people *asked* for access, not when they got it.

And that's before you even get into systems where the session itself has a start *and* an end time for the actual connection. Which one do you use then?


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

You've hit on the exact problem I ran into with ServiceNow's integration. The request `opened_at` versus the actual session `start_time` could be a full business day apart, making any "access time" report meaningless. That validation step isn't just good practice, it's the only way to know what you're actually measuring.

>Which one do you use then?

For compliance, we had to use the *later* of the session start and the credential checkout time as the official "access moment." It meant stitching three API calls together - ticket, session, and credential log - for a single audit event. The overhead was brutal, but it was the only way to get a defensible timestamp.

The real gotcha is when the session endpoint itself is rate-limited separately from the audit log. Your validation run can suddenly become a production blocker.


api first


   
ReplyQuote
(@ethan9)
Estimable Member
Joined: 3 months ago
Posts: 194
 

The checkpointing design is solid for incremental extraction, but I'd be concerned about its behavior during large historical backfills. You mention resilience to rate limiting, but the actual throttling semantics of BeyondTrust's API can be more complex than simple 429 responses. Some endpoints impose concurrent connection limits or query complexity restrictions that don't manifest as explicit rate limit headers, causing silent failures or data truncation.

Your structured JSON output is a good start, but have you considered the schema evolution problem? BeyondTrust's API fields can change between versions. If you're planning to maintain this script long-term, you'll need validation against a documented schema version and the ability to detect field deprecation. A hard-coded field mapping will break without warning.

Also, the omission of the actual timestamp field selection is a critical gap, as others have noted. The `datetime` import suggests you're handling time, but without seeing which JSON path you're extracting for the "precise time," the entire audit validity is questionable. You should explicitly document whether you're using `Request.StartTime`, `Session.ConnectionTime`, or `Session.ActivityStartTime`.


Data never lies.


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

You're focusing on pagination and checkpointing, but I'm stuck on something more basic. What library did you use for the HTTP client, and how are you handling the session/auth persistence? I tried using just `requests` with a simple session for a different API and kept hitting timeouts that weren't retrying properly because my auth token expired mid-pull. It's great it works, but that part can be sneaky.


null


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Good catch on the auth persistence. That's a common failure point when scripts run longer than the token's validity period. Even with retry logic for rate limits, a stale token will break everything.

I've seen teams handle this by either implementing a token refresh callback in their HTTP client library, or designing the script to check token expiration before starting each major batch. But then you're stuck with the question of what to do if you're halfway through a batch and it expires.

For long-running historical pulls, sometimes it's better to break the work into smaller, token-aligned chunks at the scheduler level, rather than trying to make a single script run for hours.



   
ReplyQuote
Page 1 / 4