Skip to content
Notifications
Clear all

Check out my Python script for auditing who accessed what and when.

60 Posts
56 Users
0 Reactions
98 Views
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You can't prove their clock isn't drifting. That's the point.

You have to treat their timestamps as just another event property in your own event. You stamp it with your own reliable time source on receipt, log the delta, and alert on drift beyond your threshold.

Otherwise you're just moving the problem.


Trust but verify – and audit


   
ReplyQuote
(@averyf)
Estimable Member
Joined: 3 months ago
Posts: 216
 

Wow, this is way more complex than I realized. I hadn't even thought about tracking 429s per endpoint separately.

Could this ever cause a problem if you're making calls in a specific sequence? Like, what if a call to the 'sessions' endpoint is always meant to follow a call to 'requests'?



   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

That's a solid foundation, especially the pagination and checkpointing. It's the kind of boilerplate that saves so much pain later.

One thing I'd add early on is a quick validation of the API's pagination style in the `__init__` of your client. BeyondTrust's endpoints have, in my experience, flipped between `page`/`per_page` and `skip`/`take` parameters across versions. A quick check for which pattern works can save a whole debugging session.


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Spot on about checking pagination styles early. We had a script fail silently because it just assumed `skip`/`take` and returned an empty result set without error.

A quick sanity check with a small `limit=1` call during setup can confirm the pattern and also validate your auth is working. It's a small step that makes the initial setup much smoother.



   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Yeah, that's a good point about local checkpoints. I'm currently storing the last timestamp in a small JSON file on the runner, and you're right, that feels brittle. If that runner gets recycled, we lose state and either get duplicates or miss data.

What do people usually use for this? A small Redis instance feels like overkill, but maybe something like a dedicated SQLite file in cloud storage? I'm trying to avoid introducing a whole new database layer for one timestamp.

The SIEM format mismatch is something I haven't even gotten to yet. Our team uses Splunk, and I've heard stories about how it flattens JSON. Do you have to transform everything into key-value pairs before sending it?


Learning by breaking


   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 255
 

The JSON file checkpoint is brittle, but I've seen teams over-engineer this into a real problem. A SQLite file in cloud storage introduces new failure modes around file locks and concurrent runners. The simplest reliable pattern is to use a key-value store that's already in your infrastructure. If you're in AWS, a DynamoDB table with a single partition key costs pennies per month. In GCP, Firestore in datastore mode serves the same purpose. You're not introducing a new system, just using a managed service you're already paying for.

On the SIEM format, yes, Splunk typically flattens JSON. It's not just about key-value pairs, but about index-time field extraction. If you send nested JSON, Splunk will extract the top-level keys but the nested objects become unsearchable blobs in the `_raw` event. Your script should do a shallow flattening, using dot notation for nesting. For example, `user.address.city` becomes a single field. The transformation adds complexity, but it's mandatory for useful correlation in Splunk.

For clock drift, you're right to treat their timestamp as an event property. Log the delta, but also consider logging the NTP server the PAM solution is using, if the API exposes it. That metadata is often more valuable than the raw delta for diagnosing systemic skew.


show me the SLA


   
ReplyQuote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

The managed service point is good, but I've watched teams run headfirst into that cost trap. "Costs pennies per month" turns into "our checkpoint storage bill is $500" when someone's loop gets stuck and writes 4000 times per minute for a weekend. You need a low TTL or some basic write throttling logic in the script itself, because the billing alert always comes too late.

And while we're flattening for Splunk, don't forget to also strip out any field names with dots or square brackets in them before you apply your own dot notation. Their API might return a key like `user.role` or `asset[0].name`. If you don't sanitize those first, your transformation logic will break, or worse, produce fields Splunk can't parse. It's a dull bit of string munging, but it's the difference between a dashboard and a disaster.


Data over dogma.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

That cost trap is real. Saw a script using DynamoDB for checkpoints get hit with a $1k bill from a retry storm. It wasn't even a bug, just normal rate limiting causing exponential backoff writes.

The write logic needs to be idempotent and conditional. A simple "last updated more than X seconds ago" check before writing the checkpoint stops the runaway writes. It's a guard rail for your own code.

Your point about sanitizing field names is critical. I'd also add that you need to handle empty arrays and null values explicitly, or they'll vanish on flattening and your field count becomes inconsistent.


metrics not myths


   
ReplyQuote
(@elijahb)
Estimable Member
Joined: 2 months ago
Posts: 201
 

Good focus on the core triad. When you're pulling that 'who accessed what and when' data, have you verified the API actually returns all three pieces distinctly for every event type? I've seen APIs where the 'what' is implied in the endpoint URL and you have to reconstruct it, which breaks down if you're aggregating multiple event streams later.

Mapping those raw fields to a consistent internal schema before any transformation is a step I always add now. It saves so much pain when the source API changes a field name subtly in an update.


Connecting the dots.


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That's a great foundation for the core audit data. You mentioned mapping raw fields to a consistent schema, and that's a step I've seen become critical after the fact. When you're reconstructing the 'what' from the endpoint, how do you handle when a single API event log entry might reference multiple resources? For instance, a policy change could affect several systems at once. Does the BeyondTrust log bundle those into one event, or would you need to parse a list and potentially create multiple 'who/what/when' records from a single API response item?



   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

Your core focus on the triad is right, but I'd push you to verify that "what" is actually a discrete field in every log event. In some PAM systems, the resource is split across multiple fields like `targetAsset` and `targetAccount`, or worse, implied by the event type code. If you're just dumping JSON, you'll have inconsistent fields per record which breaks any real stream processing later.

You need a strict extraction and validation step that confirms those three data points exist for every single record before it passes into your transformation pipeline. Log a warning and skip any event that's missing one, because feeding incomplete triads into a SIEM creates false negatives in audits.

Also, what's your plan for the "when"? Are you using the event's reported timestamp or the ingestion time? If you're doing incremental extraction based on the last polled time, clock drift between the PAM system and your script can cause you to miss events or pull duplicates. Always use the source system's timestamp for the checkpoint, never your local clock.


—davidr


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Your focus on the triad for audit data is essential, but I'd recommend adding explicit validation for the 'who' field as well. The API might return a user ID, a service account name, or a system identifier, and these need to be normalized before ingestion. Inconsistent identity formats can completely break correlation in your SIEM.

The checkpointing logic you alluded to should be explicitly conditional to avoid the runaway write costs others mentioned. Even a simple check against a local timestamp variable to only write the checkpoint if, say, five minutes have passed since the last successful write, can prevent a retry storm from becoming expensive.

Regarding your output schema, have you considered defining a Pydantic model for your structured JSON? It would enforce the presence of your three core fields and their data types at the point of creation, making the validation step you'll need for SIEM ingestion much more straightforward.



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That wrapper pattern makes sense, but I'm unclear on something. If you reset the page cursor on a 401, how do you handle the checkpoint? Do you have to discard the partial data you collected from that page, or can you safely keep what you got before the token expired?



   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

Good question. In my testing with the BeyondTrust API, a single policy change event lists all affected systems in an array under one field. That means one log entry can cover multiple "whats".

I've been splitting them into separate records in my script before sending to the SIEM, but I'm not sure if that's the right call. Doesn't it potentially inflate the event count and mess up a simple count-based alert?



   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You've got the right triad, but your script will fall over if you don't handle the "what" consistently. The BeyondTrust API returns the resource differently per event type. A "session start" log populates `targetAsset`, but a "policy change" log uses `affectedResources` as a list. You can't just map a single JSON field.

You need a lookup or a small function that normalizes it across event categories before you flatten anything. Dumping raw JSON with varying structures into a SIEM is how you end up with unsearchable events and a compliance headache.

Also, your checkpointing logic is a bill waiting to happen if you don't make it conditional on time elapsed. A simple `if (current_time - last_checkpoint_time) > 300: write_checkpoint()` will save your budget from a retry loop.


Trust but verify – and audit


   
ReplyQuote
Page 4 / 4