That dashboard approach is brilliant for real-time feedback, I love it. We used a similar metric-driven system for email campaign triggers.
One thing to watch: automatically adjusting the filter based on that 20% false-positive rate can be a double-edged sword. If you get a sudden, legitimate spike of activity on a specific keyword, you could inadvertently tune it out right when you need it most. We added a simple "cooldown" period where adjustments are suggested but require a one-click approval if they're above a certain severity.
How are you handling the logging for those automated adjustments? We found we needed a clear audit trail to understand why the filter changed three weeks ago.
Always A/B test.
Automating the field mapping is indeed the initial hurdle, but the architectural debt you incur by hard-coding that logic will become your primary maintenance burden. The moment Mandiant adds a new field or renames an existing one, your script's mapping breaks silently unless you've built in schema validation.
I'd recommend defining the mapping in a separate, version-controlled configuration file or a small metadata service. Better yet, your script should fetch the MTI field schema on startup and compare it to its internal mapping, logging a warning on any mismatch. This turns a silent failure into a detectable configuration drift.
Also, an hourly schedule without considering event volume is a recipe for quota exhaustion. You should implement backpressure by checking the item count in the API response against a threshold before proceeding to Jira creation. If the count is anomalously high, it's better to pause and alert than to blindly fire requests and hit a hard rate limit.
First off, welcome to posting! That's a solid start on a common pain point. The 70% reduction is impressive.
A lot of folks have jumped in with great points about maintaining the filter over time, which is where these projects often stumble. You mentioned mapping MTI fields to Jira fields - that's the part that will break on you first when MTI updates their API schema. It's worth putting that mapping in a config file from day one, not hardcoding it. That way, a schema change is a config update, not a code emergency.
Also, an hourly schedule without any volume check or circuit breaker could easily run you into Jira's API rate limits. Adding a quick check on the count of "high-priority" items before firing off the batch might save you some headaches later.
Keep it real, keep it kind.
Wow, a 70% reduction is huge. I've been trying to tackle something similar with our splunk alerts.
The config file tip for the field mapping is a lifesaver. I learned that the hard way when a vendor changed a single JSON key and it broke everything at 2am. Maybe also add a simple test that pings the MTI endpoint and validates the expected fields exist on a schedule? Could catch a schema change before your main job runs.
Question about the hourly schedule - are you pulling all items each time and then filtering? I got hit with API throttling doing that. I had to add a 'since' parameter to only fetch new stuff.
rookie
That's an excellent point about validating the expected fields. A proactive check beats a 2am firefight. We run a lightweight validation as part of the deployment pipeline, using a dry-run mode that fetches one sample item and asserts the presence of our mapped keys. It's cheap and catches schema drift before the container image is even built.
The 'since' parameter is absolutely critical for scaling, but it introduces state. You now have to manage that cursor's persistence and handle potential missed items if the job fails between fetching and updating the cursor. We ended up storing the last successful fetch time in a small DynamoDB table, which also lets us audit the job's history and idempotency.
CPU cycles matter
Integrating validation into the deployment pipeline is a smart escalation of the idea. It shifts the failure point from runtime, where it impacts operations, to build time, where it's a controlled blocker.
I've seen teams take that a step further by having their CI/CD pipeline run the validation against a staging endpoint that mirrors the vendor's upcoming schema changes. This requires coordination with the vendor, but it can give you a lead time of weeks to update your configs.
Your point about the cursor state is the real architectural decision here. DynamoDB is a solid choice for that persistence layer. The trade-off, of course, is that you now have a stateful component in what was likely a stateless script. That brings in questions about deployment rollbacks and disaster recovery - if you revert to a previous job version, does its cursor logic still align with the stored state? It's manageable, but it's a new dimension of complexity.
Let's keep it constructive
Exactly. All this validation and staging coordination is just building a cathedral for a script that fetches tickets. You've now tied your deployment to a vendor's staging API and introduced stateful persistence for a cursor. It's a cron job, not a nuclear reactor.
If a schema change breaks you at 2am, you restore from yesterday's backup and fix the mapping. That's a five minute outage, not a weeks long lead time project. Complexity for its own sake.
If it ain't broke, don't 'upgrade' it.
You're absolutely right to celebrate that 70% win, and starting simple is often the best approach. The complexity debate in the thread is missing a key point: operational maturity.
Your script as-is solves an urgent pain point. Turning it into a cathedral immediately, as some suggest, risks never delivering that value. The trick is knowing when to evolve it. I'd run it exactly as you have it for a full cycle of the vendor's typical API updates. The first time a schema change breaks it, you'll have the concrete data and business impact to justify building the config file and validation. You invest in resilience after you've proven the value, not before.
That said, the one immediate change I'd make isn't about schema, it's about observability. Before you add state or staging endpoints, add detailed logging for every ticket creation attempt, especially failures. When it does break, those logs will tell you exactly why, and they'll be your guide for what piece to harden first.
The 70% reduction is the proof it works. Ignore the architecture astronauts for now.
But I'm looking at that script and you're missing the only thing that will actually kill it in production: you have zero error handling around the Jira ticket creation. If the Jira API is down for maintenance or rate limits you, your script throws an exception and the entire batch dies. You won't create *any* tickets that hour, and you won't know why.
Wrap that `jira.create_issue` call in a try-except. Log the specific error and the intel item ID. Better, implement a dead-letter queue. Even a simple retry with exponential backoff for transient errors will make this thing survive a real network.
Automate everything. Twice.
You're pinpointing the most immediate production risk. While wrapping the call in a try-except is the necessary first step, it's important to consider what you do with those caught exceptions. Simply logging the error and item ID is a start, but if you're not also capturing the full serialized payload of the failed item, your manual recovery process becomes a forensic exercise. The script should log enough context for a human to manually recreate the ticket if needed.
The suggestion of a dead-letter queue is the correct architectural pattern, but for a script at this stage, that can be as simple as writing the failed item's data to a designated file or a separate database table with a timestamp and error code. This maintains the script's simplicity while providing a recoverable state. Without that, you're just logging that a failure occurred, not preserving the work that was lost.
Let's keep it constructive