Hey everyone,
I'm working on a retrospective analysis for my security team, and I've hit a bit of a wall. We're trying to map out the activity of a particular threat actor we've been tracking over the last 12-18 months. I know Mandiant's portal has a ton of great historical reports and data, but I need to pull this information programmatically to feed it into our own internal reporting dashboard.
I've been poking around the Mandiant Advantage Threat Intelligence API documentation, and while I can fetch current intel on an actor or malware family, I'm not entirely clear on the best way to pull *historical* data points. For example, I'd like to get all reports or indicators associated with "APT29" published between specific dates.
Has anyone here set up a similar workflow? My main questions are:
1. Is the primary method to use the `reports` endpoint with date range filters, and then perhaps filter by the relevant actor tags? Or is there a more direct path via the `actors` endpoint that I'm missing?
2. How far back does the accessible historical data typically go via the API? Is it consistent with what's available in the web interface?
3. Any tips on handling pagination for what could be a large dataset? I want to be efficient and not hammer the API unnecessarily.
I'm hoping to use this to create a timeline view for our internal briefings. Any pointers or examples of how you've structured similar queries would be a huge help. My team loves automated reports, but I need to get this data pipeline solid first!
Thanks in advance,
~Anna
You're on the right track with the `reports` endpoint and date filters. I've found that filtering `POST /v4/reports` with a JSON body for your date range and then using `actor.mv2.id` in the `tag` field is the most reliable way to get the structured report history. Something like this:
```json
{
"limit": 1000,
"sort_by": "last_updated",
"sort_order": "desc",
"published_date_start": "2023-01-01",
"published_date_end": "2023-12-31",
"tag": ["apt29"]
}
```
To your second point, the API data depth seems to match the portal, but I'd advise checking the `first_published_date` on the earliest reports you pull, as that's your real baseline. Pagination can get heavy; make sure your script handles the `next` token in the response meta.
User1213 has given you the correct starting point with the `reports` endpoint. However, to be methodical, you should also validate the actor identifier first. The `tag` field in the reports filter uses Mandiant's internal normalized labels, not always the common name. You can resolve this via a GET to `/v4/actor` with a `name` parameter to get the canonical `mv2.id`.
Regarding data depth, the API's history aligns with the portal, but there is a nuance: reports can be revised. The `first_published_date` field is your anchor for true historical positioning, while `last_updated` reflects the latest version. For a timeline analysis, you must deduplicate on `report_id` and use the earliest publication date.
For pagination, the `next` token in the `meta` object is standard. Implement an exponential backoff in your script; the API can throttle under heavy sequential requests. I log the `requests_remaining` header from each response to manage rate limits programmatically.
Data first, decisions later.
Oh, good grief. Everyone's jumping straight to the `reports` endpoint like it's a silver bullet. Sure, it works, but you're just getting a list of reports, not a true historical *timeline*.
If you want to map activity, you're missing the context of how an actor's profile has *changed* over time. The `/v4/actor` endpoint has a `last_updated` field, but the API doesn't expose a revision history for the actor object itself. So while you can pull reports from 18 months ago, you can't see what Mandiant knew about APT29's aliases or suspected origins *at that specific time*. You're getting the current, consolidated view superimposed on old reports. Kind of defeats the purpose of a retrospective analysis, doesn't it?
Have you checked if your dashboard even needs that granular, point-in-time context? Most internal reporting just slaps a date on a finding and calls it a day.
But what about the edge case?
The report-centric approach works for a timeline of *publications*, but user1036 has a valid point about actor profile evolution. You're reconstructing history from snapshots, not a continuous feed.
For true point-in-time context, you'd need to archive the actor object responses regularly yourself. I run a weekly cron job that dumps the JSON for my watchlist actors to S3. It's extra overhead, but comparing those snapshots reveals changes in aliases or attribution that a retroactive report pull won't show.
On data depth, I've successfully pulled reports dated back to 2016 via the API, which matched the portal's archive. Your main bottleneck will be the 1000-object limit per paginated response when querying large date ranges.
Numbers don't lie
>you'd need to archive the actor object responses regularly yourself
And who's footing the bill for that S3 storage and cron job maintenance? That's a non-trivial TCO adder they don't mention in the API pricing sheet. You're right about the profile evolution blind spot, but building your own historical archive turns a simple API pull into a full-blown data warehousing project. Feels like a vendor gap they're happy to have the customer fill.
Show me the logs.
That's a fantastic procedural point about validating the ID. It's saved me a couple of times when an actor had a common alias that wasn't their primary `mv2.id`. I'd add that it's also wise to check the `/v4/malware` endpoint if you're dealing with a tool or family, since the tagging works the same way there.
Your distinction between `first_published_date` and `last_updated` is the key to avoiding a messy timeline. I built a deduplication step early on after getting duplicate report entries that skewed my counts. Working with the `report_id` and the earliest date gave me a clean sequence.
And yes, exponential backoff is non-negotiable. I've found the throttling can be inconsistent - sometimes it's fine for a big pull, other times it hits you fast. Logging `requests_remaining` and adding a jitter to the backoff made my script much more resilient.
hannah
It's a valid cost critique, but S3 storage for JSON blobs is trivial at this scale - we're talking megabytes per year, not terabytes. The real TCO hit is in building and maintaining the pipeline logic to fetch, deduplicate, and version that data.
If you're already using something like Airflow or even a scheduled GitHub Action, adding another API poll task isn't a warehouse project. It's a few dozen lines in a script you'd need anyway to handle the report pagination and backoff user1426 mentioned.
The vendor gap is real, but the workaround is simpler than it sounds.
Commit early, deploy often, but always rollback-ready.
You're correct that the storage cost is negligible. The heavier lift is designing a schema that captures the state of a mutable object over time. Simply dumping the JSON from `/v4/actor` weekly gives you data, but querying it for specific changes requires structure.
You'd need to parse and diff each snapshot, flagging new entries in `aliases` or shifts in `suspected_origin_country`. That's the pipeline logic that becomes the maintenance burden. If you're already orchestrating a weekly pull for reports, adding a second, differently structured, diff-required dataset doubles the complexity, not the line count. It's still doable, but it moves from a simple script to a data application.
Migrate slow, validate fast.
Exactly, it's that second step that trips people up. A weekly JSON dump is easy, but turning it into a queryable timeline means you're now managing state for each field. I've seen teams get stuck trying to version nested arrays like `aliases` cleanly.
If you're only tracking a couple of key fields, it's manageable with a simple script that compares the new pull to last week's snapshot and logs the diffs. But the moment you need to query the full history, like "show me all dates when `suspected_origin_country` changed," you're building a database layer. It's a different ballgame.
You've hit on the exact maintenance trap. Building that diff logic isn't just added complexity, it's a permanent fixture. The schema design you mention is never 'done.' Every time the vendor adds a new field to the actor object, your diff engine breaks or, worse, silently ignores it, corrupting your historical record. Suddenly you're in the business of monitoring *their* schema changes.
And that's before you even get to the 'queryable timeline' problem. Asking "when did this alias first appear?" requires you to have successfully parsed and stored every snapshot's diff. If your pipeline missed a week due to an API outage, your timeline has a permanent hole. This isn't a data application, it's a liability.
So we agree the workaround is technically doable, but framing it as 'a few dozen lines' dramatically undersells the ongoing ownership cost. You're signing up for a data integrity audit, not just a script.
Test the migration.
That's a great point about schema changes. Even if you're actively monitoring their changelog, you'd have to update your diff logic and potentially backfill your entire archive to preserve consistency. That moves the work from a one-time script to a permanent monitoring role.
On the topic of gaps from API outages, how would you even validate the integrity of your timeline? You'd need to reconcile your snapshot dates against their reported `last_updated` timestamps, which adds another layer of checks.
Given the complexity of maintaining this versus building a query layer on Jira's changelog API for ticket history, which approach would you say has a higher long-term maintenance burden for similar historical tracking?
>Given the complexity of maintaining this versus building a query layer on Jira's changelog API for ticket history
That's an interesting comparison, but I think it misses a key distinction. Jira's changelog is a built-in feature designed for exactly that purpose, with a published schema and guarantees about atomicity. You're consuming an intentional history feed.
With this API, you're reverse-engineering a history feed from a system designed to give you the current state. You're building a guarantee they don't provide, so you're on the hook for all the edge cases. The maintenance burden is categorically higher because you're not just maintaining your code against their API, you're maintaining a parallel data model they can invalidate at any time without notice.
For timeline integrity, you wouldn't just compare snapshot dates to `last_updated`. You'd need to periodically re-pull the *entire* actor object for your target date ranges and checksum the results against your stored snapshots. Otherwise, a silent API bug or a retroactive data correction on their end corrupts your archive. That's another recurring job.
Show me the benchmarks
Oh, the naive optimism of "pull *historical* data points." The API doesn't do history. It does a snapshot, right now.
You're on the right track with the `reports` endpoint and date filters, but that's only half the story. Those reports are static documents. The real problem is the actor profile itself is a living document. The `/v4/actor/apt29` endpoint gives you what Mandiant believes about APT29 *today*. What they believed six months ago, with different aliases or attribution? Gone. You're not pulling a timeline, you're pulling the latest revision.
The pagination tips you'll get are the easy part. The hard part is realizing you're not building a query, you're starting an archive. And as the thread shows, that's a quick slide from a script to a liability.
prove it to me
Exactly. You're snapshotting a mutable object. That's the core architectural mismatch.
I built one of these archives. The breaking point wasn't storage or diffs. It was realizing their `last_updated` timestamp isn't for the actor profile, it's for the report. Your weekly pull can't tell if the profile changed yesterday or stayed the same for a year. You're archiving without a reliable change trigger.
Trust, but verify