The two-stage pipeline with a dedicated staging table is the right pattern for handling schema drift. We use a similar method, but I'd add a caveat about your versioned parsing logic.
> triggered by an `event_type` field
This assumes the `event_type` field itself is stable and reliably populated. We've seen cases where a new subtype initially lands with a null or placeholder value in that field, breaking the routing logic to your versioned parsers. Our staging process now includes a fallback check on three other metadata fields to infer the type before flagging it as unprocessable.
Your hybrid pagination method is practical, but that full scan when using the ID cursor can become expensive with very large datasets. We've implemented a cheap metadata side-table that tracks the first insertion ID for each calendar day, allowing us to approximate date filters without relying on timestamps in the API call.
Measure twice, buy once.
Yeah, that's exactly the problem. The API docs don't map to real-world use cases. You're looking for the activity log endpoint. Don't bother with the aggregated data feeds, they're useless for this.
The raw JSON is a mess, but you can get the fields you need if you request them explicitly in your query. Don't pull everything. You'll have to handle pagination and sort by insertion ID, though, because the timestamp ordering is broken.
Also, pull the asset context separately and join it yourself later. Trying to get it all in one go from the API is impossible.
—b
I feel your pain! I just went through this last week trying to build a similar dashboard for email campaign alerts. The API docs are indeed a maze.
Everyone's right about the activity log endpoint. But heads up, the rate limiting is brutal if you don't batch your requests. I learned that the hard way when my script got throttled halfway through a pull.
I'm curious, for your asset context join, are you pulling that from a separate system entirely, or is there another Anomali endpoint you're using for that data?
The rate limiting isn't just brutal, it's pathological. You have to batch, but also introduce jitter between batches. Their system seems to track concurrent sessions, so if your script gets throttled and you restart it immediately, you'll often hit the same wall. Wait a full minute, it resets.
For the asset context, there's a separate internal assets API, but its schema is completely disjoint from the activity log. You're better off pulling your asset registry from your CMDB or even Active Directory and doing a manual join after the fact. Their own join capabilities are, to be generous, fictional.
audit logs don't lie
Activity log API is the only path. Everyone's right about the ID-based pagination.
But if you're calculating latency from event generation to first action, the timestamps in the raw JSON aren't always reliable for your "generation" point. I've seen events where the system logs the ingestion time, not the original detection time. You'll need to validate that field against a known sample.
Benchmarks don't lie.