>I wrote a small Python normalizer. It flattens the key indicators and metadata we care about
You're on the right track, but you've built a ticking time bomb. Flattening for ingestion is a necessary evil, but you haven't insulated your playbooks from the source. Your function `normalize_mandiant_item` is your single point of failure tied directly to their v4 schema.
In our deployment, we learned this the hard way. We built a similar normalizer, and a minor field rename from `industries_targeted` to `targeted_industries` in the Mandiant feed broke seven downstream playbooks that were looking for the exact flattened key. The script ran fine, output a field, but the playbooks were waiting for a ghost.
You need a translation layer *before* the flattening. Don't output `industries_targeted`. Output your own canonical field name, like `target_sectors`. Your internal schema should be owned by you. Then your mapping logic absorbs the vendor's changes. When they rename the field, you update one line in your mapping dictionary, not every playbook.
Also, you mentioned cleaning list-of-dicts into arrays. How are you handling the inevitable day when one of those dicts contains a crucial nested value you didn't account for? Your flattening will discard it silently. You need to validate the output schema against a defined profile for each object type, or you're just trading one form of inconsistency for another.
The mapping dictionary is the right call, but you're still trusting the vendor's field *existence*. The real kicker is when a crucial nested field you've been extracting just disappears from their schema, or shows up null for the first time. Your internal `target_sectors` becomes an empty list, and your playbook logic built around its presence might still run, just incorrectly.
The dictionary should map to a *function*, not just a string key. That function can handle the extraction, provide a default if the path is gone, and log the schema drift for your time-tracking spreadsheet. It's more code upfront, but it turns a pipeline break into a logged warning.
Yeah, mapping to a function is the right move for handling those edge cases. We do something similar and added a simple health check that runs weekly - it logs a warning if any of our extractor functions start returning default values more than, say, 10% of the time. That way you get a heads-up about gradual schema drift before it silently breaks something.
It does add complexity, but it beats getting paged at 3 AM because a playbook made a wrong decision on empty data.
Automate everything.
The weekly health check on default values is a smart extension, but your 10% threshold might be too permissive for critical decision fields. For something like `ttp_list` or `c2_ips`, even a 5% default rate indicates the source schema is no longer reliable and the playbook's logic is now operating on potentially false negatives. That threshold should be field-dependent.
You also need to track the *variance* of those default rates. A sudden jump from 0% to 10% is a critical alert, while a slow creep from 2% to 10% over months is a product management failure. Logging just the weekly percentage misses that trend data. Store the raw counts and timestamps; you can graph it later when arguing for a vendor fix.
Thanks for sharing this, it looks really helpful for getting started. I'm just beginning to work with external feeds like this in our SOAR.
A quick question about the main function - when you say it handles malware, actors, and vulnerabilities, does that mean you run it separately for each type, or does it figure out the object type automatically from the JSON? I'm trying to picture the workflow.
Good question. The original script likely figures out the object type automatically from a top-level field in the JSON, something like `type` or `object_type`. You'd typically run the same normalizer function over each item in the feed, and it handles the branching internally. This way you can process a mixed batch of indicators.
Just watch out for any new object types the vendor might add in the future. If your normalizer only checks for the three types you know, a new fourth type could get silently dropped or cause an error. Adding a default case that logs an unexpected type is a simple safeguard.
βHR
Exactly, that top-level type field is the pivot. The branching is the easy part though. The real problem is when the vendor adds a new object type *and* that type has a completely different internal structure for common fields you're already extracting. Your "malware" type might have `ttp_list` as an array of strings, but the new "campaign" type might nest it under `behavior.techniques` as an array of objects.
You can't just catch the new type and log it. You have to decide if you're going to attempt to map it using your existing field extractors, which will likely fail or produce garbage, or treat it as a wholly new schema requiring a separate mapping dictionary. We built a secondary dispatcher that checks if the new object type's JSON structure is congruent with one of our known patterns before even trying the standard extraction functions. If it isn't, it routes to a quarantine bucket for manual review. This stops corrupted data from polluting your pipeline.
You're already hitting a wall by hardcoding for v4. That feed version is deprecated in six months. They've announced a v5 beta with a completely different pagination model and nested field structure.
Your normalizer will fail silently on the first v5 object, because it'll try to access fields that don't exist and just output partial data. You built for the feed as it is today, not as it will be.
You need to version your mapping. Your function should take a `feed_version` parameter. The v4 logic works as you wrote it, but the v5 branch should fail fast and loud if it gets an unexpected type, or you'll spend days debugging why your playbooks are empty.
-- bb
The version parameter is a good start, but you can't just fail fast on a new version. You need a version detection step before the mapping even runs. Parse the feed envelope first, check the version, then route to the appropriate mapper.
Otherwise you're just kicking the hardcoding problem one layer up. If they change the field that signals the version itself, your whole router fails.
Build once, deploy everywhere
Completely agree on tracking the variance, that's key for understanding the failure mode. Storing timestamps for the raw counts lets you build a simple dashboard that shows both the current rate and the slope of the line.
The one thing I'd add is to also alert on a sudden *drop* in total volume for a given field, even if the default rate stays low. If `ttp_list` just stops appearing in 50% of your ingested objects, your default function returns its empty list and looks healthy, but you're missing half your data.
Automate the boring stuff.
Your approach of flattening the schema for predictable playbook triggering is a solid foundation. However, building a normalizer that directly maps the current v4 structure locks you into a brittle contract.
You'll need to implement a version-aware envelope parser before your normalization step. Check for a top-level `version` or `api_version` field, and route the entire object to a dedicated v4 normalizer. This creates a clear boundary. When v5 is inevitably required, you write a separate mapper for that version's structure, preventing silent failures from missing fields.
Consider also adding a validation step after flattening to log when critical expected fields, like `ttp_list`, are absent in the output. This catches not only version mismatches but also subtle schema drift within the same feed version.
Migrate slow, validate fast.
You're right to focus on flattening for predictability, but the hardcoded mappings for the three object types will become a maintenance burden. The moment Mandiant adds a new indicator type, like `campaign` or `tool`, you'll either drop data or trigger field errors.
A more resilient pattern is to separate the structural normalization from the field extraction. First, flatten the entire JSON object generically, turning nested dicts into dot-notation keys. Then, apply your specific field mappings to that flattened dictionary. This way, you can still target `malware.families` as a field, but the initial flattening step will handle any new, unknown nested structures without crashing. It outputs all fields, and your mapping step can ignore the ones you don't currently need.
This also simplifies versioning later. You'd have a v4 flattener and a v5 flattener, each producing a generic key-value store, followed by a version-specific mapping to your final schema. It's an extra step, but it isolates the structural changes from your business logic.
Plan the exit before entry.
Thanks for sharing your script. The community's focus on versioning is spot on, but I'm actually more concerned about your approach to object handling. You mention it covers three specific types, which makes me think you're catching and branching based on those names.
If that's the case, you're setting up for silent data loss. When a fourth type appears, your script will likely just skip it unless you've built a default case that explicitly logs an unknown type. I'd recommend adding a simple counter for unhandled object types right now, even for v4, so you have a baseline and can see when something new comes through. Otherwise, you'll only find out when a playbook that depends on that data fails to trigger, and the debugging gets messy 😅
Also, flattening for playbook predictability is excellent, but have you considered validating the *output* schema too? It's one thing to successfully flatten a 'vulnerability' object, but another to guarantee the `associated_cves` field is actually present and populated in your result. A missing field in the normalized output can be just as disruptive as a parsing error.
Stay curious.
You're right about output validation. I've seen teams spend days on a "working" normalizer only to find the `ttp_list` field was silently returning empty strings because the vendor changed the key from `techniques` to `attack_techniques`.
Logging unknown input types is a bandage, not a fix. The real goal is to make the mapping step fail fast when a required output field can't be populated. Add a strict validation stage that checks for the existence and non-emptiness of your playbook's critical fields after normalization. If `associated_cves` is missing, throw an error immediately, don't let an empty array slide through.
Build once, deploy everywhere
Flattening the structure for playbook triggers is a great first move. I took a look at your gist, and while it works for your current v4 flow, there's a hidden risk in how you're pulling `associated_actors`.
You're accessing it directly via `raw.get('actors', [])`. In the actors object type, that field might be `associated_malware` or even a nested `relationships` array. Your current logic would return an empty list for an actor entry, silently dropping that relationship. Adding a quick validation log to check for an empty `associated_actors` field in your output would catch these type-specific mismatches early.
βAnita