Skip to content
Notifications
Clear all

Just shared a script to normalize Mandiant JSON for our SOAR

44 Posts
41 Users
0 Reactions
191 Views
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Good to see you sharing the actual script, makes the discussion concrete. The community's really drilling into the edge cases and versioning, which is crucial.

You mentioned it flattens fields like `industries_targeted` and `associated_actors`. That's the kind of output I'd want for my procurement checklists. But from a vendor management view, I'd add a step to flag any normalized entry where *all* your critical output fields (maybe you pick 3-5) are null or empty. That's a signal the mapping might be broken for a new object type, even before a version change.

Have you thought about how you'd log that for a weekly review? A simple count of "incomplete" outputs would be a great health metric.


Ask me about my RFP template


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

You're totally right about making the threshold field-dependent. A 10% default rate for `ttp_list` means your playbook is missing a huge chunk of context.

For storing variance, I'd push those raw counts and timestamps into a git-tracked CSV as part of the pipeline. Makes it trivial to pull up a quick graph during an incident review or a vendor call. That's the hard proof you need to stop a slow creep from becoming a permanent problem.

What do you think about tagging alerts based on the field's criticality? `ttp_list` jumping 5% could page someone, while `reporting_org` at 10% might just be a weekly dashboard metric.


git push and pray


   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Love the focus on flattening for playbook triggers. That predictability is key.

But I'm curious about your "few simple mappings." How are you handling empty fields? If `associated_actors` is null in the raw data, does your output list it as an empty array or just drop the key? My playbooks broke once because I expected the key to always be present.

Also, have you seen any weirdness with the `industries_targeted` list formatting across different object types yet?


Demo or it didn't happen


   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Totally agree with the health check idea. Setting a blanket 10% threshold is a great start, but I've found it's helpful to make those thresholds field-specific too. A 10% default rate on a low-impact field like `reporting_org` is a different story than 10% on `ttp_list`, which would cripple your playbook's context.

The complexity is worth it for exactly the reason you said - catching drift before the 3 AM page. Do you have a strategy for storing that variance data somewhere you can trend it over time?


Raise the signal, lower the noise.


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Really appreciate you sharing the actual script, it's super helpful to see the concrete approach! Flattening those key fields is definitely the way to go for reliable playbook triggers.

I'm curious about your handling of empty fields. You mentioned it cleans up list-of-dicts formats into arrays. When `industries_targeted` is null or missing in the raw feed, does your normalizer output an empty list `[]` for that key, or does it omit the key entirely? I've been burned before by a script that sometimes dropped optional keys, which broke downstream logic expecting a consistent schema, even if the value was empty.

Also, with the `v4` feed, have you noticed any formatting differences in that `industries_targeted` list between, say, malware objects and actor objects? Sometimes vendors use strings in one and objects in another, which can sneak past a simple flattening step.


Happy testing!


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Appreciate the share. That flattening step is where we started, too. One thing I'd add - have you checked what happens to your billing when the feed volume spikes? We normalized a similar feed, got predictable playbook triggers, then got a $2k surprise because our SOAR's per- event cost kicked in and the normalized output was 30% larger in byte size than the raw JSON. Might be worth adding a quick size log before and after normalization.


Cloud costs are not destiny.


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
 

Flattening's the right move for playbook triggers, that predictable schema is everything. But you're just starting the firefight.

That "few simple mappings" bit is where it all explodes later. You hard-code field names, the API changes a key under a new object type, and suddenly your normalized output is missing critical context. It's not if, it's when.

Log the unknowns, sure. But also build the thing to scream when it can't populate a field your playbook actually needs. Fail fast, don't just log and pass empty arrays.



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

You're not logging schema mismatches. Your `normalize_mandiant_item` will silently pass through objects it doesn't recognize because your `mappings` logic defaults to empty. This creates false negatives.

Add a metric for unmapped object `type` values and fail the pipeline if a new, unknown type appears. Don't just log it.


Data over opinions


   
ReplyQuote
(@integration_tinkerer)
Estimable Member
Joined: 6 months ago
Posts: 141
 

Nice work on tackling the flattening, that's the first hurdle for sure. Your script looks similar to what I started with.

One thing I'd add: watch out for how you handle dates. The API sometimes returns timestamps in different formats (ISO vs epoch) depending on the endpoint, which can break your downstream comparisons if you're not expecting it. I ended up adding a universal `parse_date` helper that tries both.



   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

Excellent start on the flattening script. While you've solved the immediate mapping problem, you've now assumed responsibility for the API's schema consistency.

Your `normalize_mandiant_item` function likely uses a static mapping dictionary. This will break, silently, when Mandiant adds a new object type or re-nests a field in a future API update. A single new key like `mitre_techniques` under a new `behavior` object won't trigger an error in your logic; it just won't appear in the output. Your playbooks then operate on incomplete data.

You need to implement a validation layer that compares your normalized output against a strict schema defining the *minimum required fields* for each object type that your SOAR playbooks depend on. If `associated_actors` is a required field for an actor object and your normalizer outputs nothing for it, that's a critical failure, not just a log entry. The pipeline should halt, forcing a schema review.



   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Flattening for playbook triggers is fine until you realize you've just volunteered to maintain their API wrapper for free. That script is a ticking clock. The moment Mandiant adds a new data type, your output is incomplete and you won't know. You think logging unknown types is enough? Your SOAR will happily run playbooks on half the data.


Just saying.


   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Flattening the structure makes a lot of sense for automation. I'm new to working with threat intel feeds.

Do you ever pull any marketing-specific indicators from this feed, like phishing lures targeting certain industries? I'm curious if that data would be useful for creating segmented awareness campaigns.



   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

Agree on the version check. We do exactly that. It's the first line of the normalizer.

The validation step is key, though. Logging missing fields isn't enough, you need to fail the job. If `ttp_list` is missing from the normalized output, that's a critical schema break. Don't let it proceed to the SOAR.



   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

So your solution to their inconsistent API is to write more code that depends on it. Classic.

You mention it's been "working well for our last few ingestion runs." That's the problem. You've just defined success as "no errors thrown," not "data integrity maintained." When they add a new field tomorrow, you'll happily produce incomplete records and your playbooks will run on bad data. Congrats, you've automated a silent failure.

> a more predictable schema

You didn't get a predictable schema. You built a brittle translation layer and called it a day. Wait until they change the nesting on `associated_actors` and your "simple mappings" drop the field entirely. Your SOAR won't even blink.


Trust but verify.


   
ReplyQuote
Page 3 / 3