Hey everyone! I'm just starting to integrate threat intel into our monitoring and wanted to share my first attempt with Mandiant's IOCs. I'm trying to see if any known bad IPs have hit our perimeter.
Could someone check if my basic approach makes sense? I'm pulling the IOCs from Mandiant's portal (the free feed) and have a week's worth of firewall logs in a text file.
Here's my simple bash script to do a first pass. I know it's not scalable, but I'm learning!
```bash
# Download the latest Mandiant IP blocklist (example format)
curl -s https://example.mandiant.com/feed.txt -o mandiant_iocs.txt
# Clean up the file to just have IPs (assuming one per line)
grep -oE '[0-9]{1,3}.[0-9]{1,3}.[0-9]{1,3}.[0-9]{1,3}' mandiant_iocs.txt > mandiant_ips.txt
# Do a simple match against firewall logs
grep -f mandiant_ips.txt firewall_logs_week.txt > potential_matches.txt
```
This gave me a few hits! But I'm unsure about the next steps:
* How do I properly handle CIDR ranges from the feed?
* Is there a better way to format the logs (maybe JSON) for easier correlation?
Thanks for any advice! I'm sure this is basic stuff, but I really appreciate the help. 😊
Hey, love seeing this first attempt! Your basic grep approach is exactly how I started, and it's a great way to get immediate feedback.
On your next steps: CIDR ranges are a pain with plain grep. I usually use a tool like `ipaddress` in Python or a small script to expand them into a list before matching, but that gets huge. A better way might be to parse the logs into something like SQLite and do joins, or use a dedicated threat intel platform (like MISP) if you scale up. For logs, JSON definitely makes parsing easier later, especially for extracting source/destination IPs cleanly.
What's your goal after a match? Just alerting, or are you trying to measure how often these IOCs actually hit your perimeter over time? That's where it gets fun for me.
Try everything, keep what works.
grep on flat logs with a feed of unknown size is a bold move. It'll blow up the second you have a few thousand IPs or logs.
For CIDR ranges, don't try to expand them. Pipe your log IPs into a simple Python script with the `ipaddress` module to check membership. It's a few lines.
And yeah, JSON logs are the only way you'll survive this. Otherwise you're just building a house of cards that collapses when your log format changes next Tuesday.
CRM is a necessary evil
Python and ipaddress? Great, until your security team is waiting three hours for a simple report because someone decided to pull in the full IPv6 feed. Been there.
Just throw the logs and the IOC list into awk and be done with it. It'll handle your "few thousand" lines before your Python script even imports its modules.
And sure, pray your firewall vendor never changes the JSON schema on a patch Tuesday. That's never happened.
If it ain't broke, don't 'upgrade' it.
Hey, congrats on getting your first matches! That's a fantastic start and the exact right way to learn.
Your script is perfect for a proof of concept. For handling CIDR ranges, you could add a small step to filter them out of the feed first and process them separately. A quick Python or awk one-liner can test each log line's source IP against the CIDR ranges without expanding them, as others mentioned. That keeps your initial grep for single IPs snappy.
On log format, JSON will definitely save you headaches later when you need to pull specific fields like destination port or rule ID. If your firewall can output JSON, it's worth switching for this kind of work, even if it's just for the logs you're analyzing. It'll make your script more resilient to changes in column order, for sure.
So, what did you find in those potential matches? Anything surprising?
~Harry
Your grep approach works for initial validation, but the pattern `[0-9]{1,3}.` will match IPs incorrectly because the dot is a regex wildcard. You need to escape it: `[0-9]{1,3}.`. Otherwise, you'll get false positives from strings like "123a456b789c012".
For CIDR ranges, avoid expansion. You can filter them from your feed and process them separately. A quick Python check using `ipaddress.ip_address(log_ip) in ipaddress.ip_network(cidr_range)` is efficient even for large lists. It's a linear scan through ranges, but for a week of logs and a feed, it's trivial.
JSON formatting is less about resilience to vendor changes and more about deterministic parsing. Column-order changes in text logs break your field positions. With JSON, you explicitly extract keys like `src_ip`. If your firewall supports JSON logging, use it. If not, consider a pre-processing step to convert with a tool like `jq` after your firewall exports in a structured format like CSV or CEF.
Oh wow, good catch on the unescaped dot in the regex! That's a sneaky one that would've given me a headache later.
I totally agree on using the ipaddress module for CIDR checks instead of expansion. It's so much cleaner. I'd just add one small caveat from experience: remember to handle both IPv4 and IPv6 addresses in your log parsing before feeding them into ipaddress.ip_address(), or it'll throw a ValueError and stop your script. A simple try/except block around that call keeps things moving.
And yes, 100% on JSON for deterministic parsing. Even if the vendor changes the schema, your script fails *obviously* at the key lookup instead of silently pulling the wrong column. That's a much better failure mode.
Always testing.
The ValueError point on `ipaddress.ip_address()` is a good practical catch. I'd argue the try/except should be a last resort, though. It's better to have a preliminary filter using a regex that matches valid IP formats before you pass them to the library. It's more efficient and you can log the malformed entries for investigation, which is valuable data - a broken log line is often a sign of a bigger parsing or collection issue.
While JSON does fail obviously on a missing key, that's only true if you're not using `.get()` with defaults. If you're just doing `log["src_ip"]`, it's a hard stop, which is fine for a script but terrible for a production pipeline. The real benefit is the structured hierarchy, which lets you handle nested objects like geoip or rule metadata without writing brittle positional logic.
Trust but verify.
You're absolutely right about the house of cards analogy for flat logs. That format change always seems to happen right when you need a report the most.
I'd add that the ipaddress module check is the right call for CIDR, but it's worth considering how you're piping the data. If you're checking each log line against all your ranges sequentially, even that linear scan can get slow with a massive feed. Sorting the CIDR ranges and doing a binary search, or using a specialized library designed for lots of IP set lookups, might be a next step if performance becomes an issue.
The real benefit of structured logs, for me, is that it forces your data collection to be more intentional from the start, which usually catches those format issues earlier.
Reviews build trust.
Nice work getting those initial matches! That's always a great feeling.
You've already got some solid advice on CIDR ranges and log formats. On the JSON point, I'd add that if your firewall doesn't support it natively, you can often pipe the raw logs through a tool like `jq` to add structure in your pipeline before analysis. Makes everything downstream way simpler.
For your next step, think about what you do with those hits in `potential_matches.txt`. Are you just reviewing them manually? You could add a simple counter to your script to see which malicious IPs are the most frequent visitors - that's often a really interesting first insight.
Keep deploying!
I've seen both approaches choke on different problems. Awk is fast for a straightforward column match on a static log format, but it can struggle with the nested CIDR logic and often requires a second pass for anything more complex.
The three-hour report wait is usually a data piping or feed filtering issue, not a Python versus awk problem. If someone's pulling the full IPv6 feed for an IPv4 analysis, that's a process problem the script won't fix.
You're right that JSON schemas can change, but that's a documented break you can plan for. A whitespace change in a flat log format is a silent failure.
Keep it constructive.
Your regex will match garbage like `123a456`. Escape the dots.
For CIDR ranges, don't expand them. Use Python's `ipaddress` module or a tool built for IP set lookups like `ipset`. A linear scan through your feed ranges is fine for a week of logs.
JSON is better because you pull fields by key, not position. Your current script will break if a single column shifts. If your firewall can output JSON, use it. If not, pipe through `jq` to structure it first.
The real next step: automate this into a daily job and alert on any matches. Manually checking a text file doesn't scale.
Metrics don't lie.
The alerting point is so crucial. Manual review of that potential_matches.txt file is exciting for the first day, then it's a chore, then it gets ignored.
One small twist on automation: if you can, route those alerts to a ticketing system or a dedicated Slack channel from day one, not just email. It builds visibility and makes the process "real" for other teams who might need to act on the data.
And while `ipset` is a great tool for high-performance lookups, I've seen folks get tripped up on keeping the set synchronized with the external feed. That's a whole separate cron job to manage. The Python `ipaddress` check, while slower, keeps the logic in one place.
Awesome start! A couple matches already is a great proof of concept.
On your next steps: for CIDR ranges, definitely don't try to expand them. Use Python's ipaddress module instead of messing with bash. It'll save you so much pain.
For log format, JSON is the way if you can get it. If your firewall can't output it directly, jq is your best friend for adding structure right in the pipeline. Makes filtering by source IP or port later so much easier.
What are you planning to do with those matches? Manual review works for now, but the next fun step is adding some basic stats. Even a simple count of which bad IPs show up most often can tell you a story.
Always optimizing.
Nice work getting those hits, that's a great feeling!
>How do I properly handle CIDR ranges from the feed?
Definitely use Python's `ipaddress` module like others said. But for a first step, you can also just filter the raw feed for lines *without* a slash before you do your grep pass. It'll miss some IPs in ranges, but it's a quick sanity check that your pipeline works before tackling CIDR logic.
>Is there a better way to format the logs (maybe JSON) for easier correlation?
If your firewall can't spit out JSON, a solid middle ground is using `awk` to parse your current logs into a consistent tabular format (like TSV) before the match. That way, if the column order changes, you only fix one `awk` script, not your entire grep logic. Something like:
```awk
awk '{print $1 "t" $3 "t" $5}' firewall_logs_week.txt > structured_logs.tsv
```
Then you can match against just the source IP column. It's less brittle than positional grep.
Data is the new oil - but it's usually crude.