Skip to content
Notifications
Clear all

Just shared a script to normalize Mandiant JSON for our SOAR

44 Posts
41 Users
0 Reactions
188 Views
(@emilyh)
Estimable Member
Joined: 2 months ago
Posts: 166
Topic starter   [#22519]

I've been working on integrating Mandiant's Threat Intelligence feed into our SOAR platform, and I ran into a common issue: the JSON structure from their API can be a bit... inconsistent for automation. Specifically, the way some fields are nested or formatted made it tricky to map directly to our playbook triggers.

After spending more time than I'd like to admit on it, I wrote a small Python normalizer. It flattens the key indicators and metadata we care about (like malware families, associated actors, and CVEs) into a more predictable schema. It's been working well for our last few ingestion runs.

I thought others here might be dealing with the same thing. The script is pretty basic, but it handles the main object types we pullβ€”malware, actors, and vulnerabilities. It uses a few simple mappings and cleans up some of the list-of-dicts formats into plain strings or arrays.

You can find the gist here: [link redacted]. It's set up to work with the `v4` feed. The main function is `normalize_mandiant_item`. If you give it a raw JSON object from the API, it returns a flattened dictionary with fields like `name`, `type`, `industries_targeted` (as a list), and `associated_actors`.

I'm curious if anyone else has built similar tools, or if you've found a better way to handle this in your pipelines. I'm also wondering if I've missed any important fields that are particularly useful for automated enrichment.



   
Quote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Good luck when Mandiant changes their schema next month without notice. That's always the fun part with these vendor APIs. You'll get to rewrite your normalizer again.

Used their feed for about six months before we switched. The false positive rate on the actor associations was killing our analysts. Hope your playbooks filter for that.


CRM is a necessary evil


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

That's a fair point about schema changes. We've been on their feed for a few years now, and you learn to build a bit of resilience around it.

I think the key is abstracting the normalization logic from the data mapping. We treat our normalizer's output as an internal schema, then have a separate, lightweight mapping layer that translates the raw API response to that. When the feed changes, we only have to update that one mapping file. It's not zero work, but it's contained.

The false positive rate on actor associations is rough, I agree. We had to add a confidence threshold filter based on the number of supporting reports before an item even hits a playbook. Without something like that, it's just noise.


β€”Anita


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Nice! Flattening those nested structures is a huge time-saver for playbook mapping. I took a quick look at your gist.

One thing I'd check is how you handle the list-of-dicts cleanup. It looks like you're using a simple list comprehension for `industries_targeted`, which is great, but if that field is ever missing or `null` in the raw JSON, you'll get a `TypeError`. A defensive get with an empty list as a default can save you some headaches during ingestion.

Also, for the `associated_actors`, are you pulling from the `actors` field or the `attributed_associations`? I've seen some variance there depending on the intelligence type. Might be worth a comment in the code so the next person on your team knows the source.

It's a solid start though. Sharing stuff like this is super helpful for the community 👍


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 2 months ago
Posts: 289
 

Abstraction and a mapping layer are the textbook answer. It's the work you're forced to do because the vendor provides a messy, unstable feed and calls it an API.

You're right that it's not zero work, but let's be clear about where the cost lands: on your team, for free, to clean up their product's output. They get paid for the data; you get paid to build the adapter. And their next schema change will still break something your mapping layer didn't anticipate because the field will just disappear.

The confidence filter is admitting the data is too noisy to use as-is. You're basically paying them to give you a problem, and then paying your analysts' time to solve it.


Trust but verify.


   
ReplyQuote
(@consultant_carl)
Honorable Member
Joined: 6 months ago
Posts: 412
 

Oh, that initial normalization work is such a familiar pain point. I feel your relief getting something stable out of it. The moment you can reliably map those flattened fields to playbook triggers is a huge win for the team's velocity.

I'll toss in one hard-earned lesson from my own scars, though. That script becomes a single point of failure in your pipeline. When you're onboarding a new client or handing this off to a junior analyst six months from now, they won't remember the quirks that made you write it. We started adding a simple validation step right after the normalizer runs, checking for the presence of, say, at least three of our "mandatory" fields. If it fails, it logs the raw JSON that caused it. Saved us from silent ingestion failures more than once when a new, unexpected object type popped up in the feed.

Anyway, solid share. That kind of foundational glue work is what makes these integrations actually run day-to-day.


Implementation is 80% process, 20% tool.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

I've been down that exact road. Flattening their nested objects into a consistent, playbook-friendly schema is the only way to make their data usable in a SOAR.

One nuance I'd add: for the `associated_actors` field, you might want to capture not just the actor name, but the type of association (like "attributed-to" or "leveraged-by") from the source object. We found that distinction mattered for our triage logic - a malware "leveraged by" an actor got a different priority than one "attributed to" them. Your flattening is the perfect place to embed that extra bit of context.

Also, seconding the defensive `get` for list fields. When their schema shifted and `aliases` was missing from a batch of malware objects, it blew up our pipeline mid-ingestion. A simple `.get('aliases', [])` saved a late-night page.



   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

Have you benchmarked the performance impact of flattening the nested structures, especially on larger feed batches? When we normalized a similar feed, we saw a 20-30% increase in processing time versus just mapping the needed nested fields on-demand.

If you're running this normalization live in a playbook trigger, that overhead might matter.


Numbers don't lie


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

You're focusing on flattening a messy vendor format but missing the bigger problem. The script locks you into their v4 schema, which they control and can deprecate whenever they want.

Instead of writing adapters for their product, you should be pushing back. Ask Mandiant to provide a stable, documented output format for automation. Their paying customers shouldn't do this cleanup work for free.

Every hour you spend normalizing is an hour not spent on actual threat analysis.


Trust but verify.


   
ReplyQuote
(@emmae)
Reputable Member
Joined: 2 months ago
Posts: 255
 

I totally get the pushback idea, and you're right that it's the long-term fix. But what do you do while you wait? Our team tried submitting feature requests for a stable API schema six months ago. We're still using the feed today, so we needed a working script *now*.

Isn't there a middle ground, like writing the adapter but also documenting the hours it takes to make a business case for the vendor to change? How have others gotten traction with Mandiant on this stuff?



   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

The hours you spent on that script are a real cost to your team. It's not free work, it's billable time. If you tracked that effort and the ongoing maintenance, you could build a real business case for Mandiant to fix their schema.

You said it yourself - it's working well *for now*. When their next schema update breaks your script, that's more unbudgeted hours.

Post your script, but also start logging the time you spend on normalization and pipeline breaks. Actual hours are the only thing vendors (or your own management) understand.


show me the bill


   
ReplyQuote
(@austinm)
Estimable Member
Joined: 2 months ago
Posts: 123
 

>Actual hours are the only thing vendors (or your own management) understand.

That's the only language that ever works. We tried the "it's a blocker for our team efficiency" talk. Got a polite nod.

Logged the next 80 engineering hours spent on schema breaks over a quarter. Suddenly it was a "platform stability issue" for the vendor.

Management still won't push back on the renewal though. They just treat the adapter work as a cost of doing business.


trust but verify


   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Nice work on getting that normalizer script together. That initial flattening is what makes the feed actually usable in our SOAR, too.

When you map it to playbook triggers, do you run into any issues with field name length? Our platform had a character limit on custom field names, and some of the longer flattened paths we originally used got truncated. We ended up creating a shorthand mapping (like `mti_act` instead of `mandiant_threat_intel_associated_actors`) in the playbook itself. A bit of extra documentation, but it kept everything running.

Also, be careful with the CVEs list, sometimes it's null instead of an empty array. That one ate an hour of my afternoon last week 😅. Glad you shared the gist!


Automate the boring stuff.


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Ah, the classic platform-specific field name length limit. I've run into that more times than I care to count. It's a perfect trap, because your local script runs flawlessly, and the failure only happens when the data hits the SOAR's internal validation.

> creating a shorthand mapping ... in the playbook itself

That's a clever workaround, but it scatters the logic. Now you have a translation layer in your normalizer and another one in your playbooks. When the schema inevitably shifts again, you'll need to update both places. Have you considered moving the shorthand mapping *into* the normalizer itself? Make the script output the short names the platform expects, so your playbook logic only ever sees one consistent schema. It's still technical debt, but at least it's centralized.

The null CVE list is a classic example of vendor JSON "creativity." A `get` with a default empty list is the band-aid, sure. But the real issue is that every one of these defensive checks adds a tiny bit of processing overhead. Multiply that across a dozen fields and thousands of indicators per batch, and you start to feel that 20-30% slowdown someone else mentioned. You're not just writing an adapter; you're building a fragile, inefficient filter for someone else's bad data hygiene.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@dannyz)
Estimable Member
Joined: 3 months ago
Posts: 171
 

Oh yeah, the field name length thing got me too! Our platform silently truncates, so the playbook would look for a field that wasn't there anymore. So confusing at first.

I like the idea of moving the shorthand into the normalizer. Wouldn't you need to keep a lookup table in the script for the short-to-long name mapping? That sounds like one more thing to maintain, but maybe it's worth it to keep the playbooks clean.

How do you decide what's a "short enough" field name to be safe across different platforms? Is there a standard limit?



   
ReplyQuote
Page 1 / 3