Hey folks, data_shipper_joe here 👋
Just wrapped up a project to get our Ping Identity logs flowing into our SIEM (Splunk, in our case). We needed better visibility into auth events for security monitoring, and while Ping offers great APIs, getting that data streaming reliably was a bit of a journey. Since I live in the data integration world, I figured I'd share the approach.
We ended up building a lightweight service using Python and the PingDirectory REST API. The key was using the `/log-files/access` endpoint with careful filtering and pagination to avoid missing events. We run this as a container on a schedule, checkpointing the last read timestamp to avoid duplicates.
Here's the core of the fetch logic:
```python
import requests
import time
def fetch_ping_logs(last_timestamp):
url = "https://your-ping-server:1443/log-files/access"
params = {
'filter': f'(timestamp>={last_timestamp})',
'sortOrder': 'ascending',
'pageSize': 1000
}
headers = {'Authorization': 'Bearer YOUR_SERVICE_ACCOUNT_TOKEN'}
response = requests.get(url, params=params, headers=headers, verify=False)
logs = response.json()['logs']
# Process logs, send to SIEM HTTP Event Collector...
new_last_timestamp = logs[-1]['timestamp'] if logs else last_timestamp
return new_last_timestamp
```
The main "gotcha" was handling the SSL certificates properly in our container and managing the rate limitsβPing's API can be strict. We also had to map their log schema to our SIEM's CIM model for consistency.
Anyone else done something similar? Curious about alternative methods, maybe using a syslog forwarder from Ping directly? Also happy to help if anyone's tackling a similar data-shipping project with auth logs.
ship it
ship it
You've essentially built a polling service with state management, which is the obvious way to do it when a proper streaming API isn't on the table. I've seen teams spin up entire Kubernetes deployments with Helm charts for this same pattern, which is frankly overkill.
One thing I'd check is how that `verify=False` is sitting in production. That's a shortcut that'll make any security team's SIEM alerts go off, ironically enough. At the very least, pin the certificate.
Also, what's your schedule interval? If it's too aggressive you'll hammer the Ping API and likely get throttled, too slow and you'll have a visibility gap. I've found that most log ingestion of this type is fine with a 5-minute poll, but you need to size your pageSize to handle the potential burst.
keep it simple
Totally agree about the certificate pinning being crucial. Even for an internal service, that `verify=False` can cause issues with some corporate proxies or future security scans.
On the polling interval, 5 minutes is a good default. One thing I've run into is that the "right" pageSize isn't just about volume, it also depends on the API's max limit and how it handles sorting. If you set it too high and the sort is unstable near the cutoff, you might miss an event during a busy period. I usually start with something like 500 and test with a simulated spike.
Also, have you seen any issues with the log endpoint's response time varying? That's another reason to be conservative with the schedule.
Clean code, happy life
The `verify=False` is a red flag for cost as much as for security. If your security team catches that and forces a re-architecture mid-stream, you're burning engineering hours that could have been avoided by pinning the cert upfront. I'd argue the cost of doing it right the first time is lower than the ops overhead of a ticket and a redeploy.
On the polling interval: 5 minutes works, but you should also factor in the cost of the API calls themselves. If you're on a usage-based pricing tier with Ping, each request has a marginal cost. A 5-minute poll with a pageSize of 500 and a burst of 2000 events per poll means you're making 4 API calls per cycle. That's 1,152 calls per day. If your Splunk ingest pricing is also volume-based, you might want to model the total cost per event. I've seen teams set the interval to 10 minutes without any noticeable visibility gap and cut their API consumption in half.
What's your actual event volume per day? And are you using any caching or dedup on the Splunk side to avoid double-counting if a poll picks up a previously checkpointed event?
Spreadsheets or it didn't happen.
`verify=False` is a classic shortcut that never dies. It'll bite you during an audit or when someone decides to run a vulnerability scan. You'll spend more time justifying that line than you did writing the rest of the service.
Also, hope your `last_timestamp` logic is airtight with that ascending sort. If your checkpoint fails mid-page, you'll get duplicates or, worse, miss logs entirely. Seen that happen more than I'd like.
SQL is enough
That checkpoint failure scenario is a perfect example of why I push for atomic operations in these log shippers. If your state save isn't part of the same transaction as your successful log batch send, you're building a time bomb.
I've seen teams try to solve it with increasingly complex idempotency logic, when the simpler fix is often to checkpoint *before* you fetch the next page. You process from the saved timestamp, but you only advance the saved timestamp after you've confirmed the entire batch is in the SIEM. Means you might reprocess a few logs after a crash, but your SIEM's deduplication should handle that, and you never lose data.
keep it simple