Skip to content
Notifications
Clear all

Check out my script to pull daily high-confidence IOCs into our SIEM.

2 Posts
2 Users
0 Reactions
9 Views
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
Topic starter   [#27095]

Having evaluated numerous threat intelligence platforms for operational integration, I've consistently found the most critical gap to be the automated, reliable, and credentialed ingestion of high-confidence indicators into core security telemetry systems. Many vendors provide feeds, but the burden of filtering, deduplication, and secure delivery often falls to the consumer, creating fragile scripts and alert fatigue.

To address this within our ThreatConnect deployment, I've developed and productionized a Python-based orchestrator that executes daily, pulling only IOCs with a confidence rating of 75 or above and a threat rating of "High" or "Critical" from specified sources and groups. The primary design goals were idempotence, comprehensive logging for audit, and direct integration with our SIEM's HTTP Event Collector (HEC). The script performs the following key operations:

* Authenticates to the ThreatConnect API using TC_API_KEY/TC_SECRET_KEY stored in a secrets manager.
* Applies temporal and confidence filters to the indicator query to limit scope to newly added or modified items from the last 24 hours.
* Performs local deduplication against a small SQLite state table to prevent re-sending identical IOCs.
* Transforms the relevant IOC fields (type, value, rating, confidence, description, source) into a JSON schema normalized for our SIEM.
* Batches and forwards the events via a mutually authenticated TLS session to the SIEM ingress point.
* Logs all operations, including counts and any failures, to both stdout (for container logs) and a dedicated audit index.

The core of the retrieval logic is as follows:

```python
def fetch_high_confidence_iocs(tc, from_date):
"""Fetch IOCs with confidence >= 75 modified since from_date."""
indicator_filter = {
'confidence': ('>=', 75),
'lastModified': ('>=', from_date),
'threatRating': ('IN', ['High', 'Critical'])
}
try:
indicators = tc.indicators()
indicators.set_filter(indicator_filter)
# Limit fields to reduce payload
indicators.set_fields(['type', 'summary', 'rating', 'confidence', 'threatAssessScore', 'dateAdded', 'lastModified', 'source', 'description'])
return indicators.retrieve()
except RuntimeError as e:
logger.error(f"TC API retrieval failed: {e}")
raise
```

**Performance & Cost Observations:**
After 90 days of execution in a Kubernetes CronJob (allocated 200Mi memory, 0.2 CPU), the script processes an average of 1200 indicators per run. The runtime averages 45 seconds. This translates to negligible cloud compute cost (~$0.02/month). More importantly, by shifting the filtering and transformation logic upstream of the SIEM, we've reduced our ingest-related licensing costs by an estimated 8-10% compared to pulling a full unfiltered feed.

**Key Pitfalls to Avoid:**
* **API Pagination:** The default return limit is 500. Implement a loop to handle `next` tokens for large result sets.
* **State Management:** The SQLite state table must be mounted from a persistent volume in a containerized environment to avoid resending the entire history.
* **Credential Rotation:** Integrate with your platform's secret rotation lifecycle. The script fails gracefully if the API keys are invalid.
* **Schema Drift:** Validate the JSON output against your SIEM's expected schema periodically, as ThreatConnect field mappings can change during platform upgrades.

This approach has provided us with a deterministic, maintainable pipeline. I am interested in critiques of the methodology, particularly regarding the confidence threshold heuristic or alternative state management strategies beyond a local SQLite file. Has anyone conducted a comparative analysis of the correlation efficacy of IOCs filtered at different confidence levels versus the signal-to-noise ratio in their alert queues?

-ek


Show me the numbers, not the roadmap.


   
Quote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

The confidence threshold filter you've implemented is a solid starting point, but have you considered its interaction with the temporal filter? In my experience, an IOC's confidence score can be revised upward days after its initial publication, often following broader community validation. A strict "last 24 hours" window might miss these critical updates.

You might add a secondary check for any indicator, regardless of age, whose confidence rating has been elevated to meet your 75+ threshold since the last run. This catches those slower-moving, high-fidelity indicators that mature after the initial reporting cycle.

What's your strategy for managing the lifecycle of these indicators in the SIEM once they're ingested? Automated expiration based on the IOC's own last-seen or updated timestamp is something I've seen teams forget to build in.


Measure twice, spend once


   
ReplyQuote