Our Security Operations Center recently faced an issue where internal SIEM alerts for outbound connections to known-bad IPs lacked crucial context. We could see the connection attempt, but we had no immediate insight into the threat actor's affiliation, the associated campaign, or the confidence in the intelligence. This led to a manual, time-consuming process of pivoting to the Recorded Future portal. To streamline this, I designed an integration to correlate our internal event logs with Recorded Future's intelligence directly within our SIEM (Splunk, in this case). The goal was to enrich our alerts in real-time, providing analysts with a consolidated view.
The architecture hinges on two primary Recorded Future APIs and a middleware service I built to handle transformation and rate-limiting. The core components are:
1. **Internal Event Feed:** Our firewalls and proxies send connection logs to Splunk, where we identify events of interest (e.g., connections to IPs on our internal watchlist).
2. **Middleware Enrichment Service:** A lightweight Python service subscribes to a Splunk HTTP alert webhook for these events. For each event, it performs the following steps:
* Extracts the observable (e.g., destination IP).
* Calls the Recorded Future `Connect API` for that IP entity to fetch the risk score and associated evidence.
* If the risk score exceeds a defined threshold, it makes a secondary call to the `Context API` to obtain detailed threat intelligence (e.g., malware families, related campaigns).
* Formats this enriched data and sends it back to Splunk as a new event, linked to the original via a correlation ID.
Here is a simplified excerpt of the enrichment service logic, showing the key API interactions:
```python
import requests
def enrich_ip_with_rf(ip_address, rf_api_token):
# Step 1: Fetch risk and evidence via Connect API
connect_url = f"https://api.recordedfuture.com/v2/ip/{ip_address}"
headers = {'X-RFToken': rf_api_token}
params = {'fields': 'risk,evidenceDetails'}
connect_response = requests.get(connect_url, headers=headers, params=params)
connect_data = connect_response.json()
enrichment_payload = {
'original_ip': ip_address,
'rf_risk_score': connect_data.get('risk', {}).get('score'),
'rf_risk_reasons': [e.get('rule') for e in connect_data.get('evidenceDetails', [])]
}
# Step 2: If high risk, get context
if enrichment_payload['rf_risk_score'] > 80:
context_url = f"https://api.recordedfuture.com/v2/ip/{ip_address}/context"
context_response = requests.get(context_url, headers=headers)
context_data = context_response.json()
enrichment_payload['rf_context'] = {
'intel': context_data.get('intelCard'),
'related_malware': context_data.get('relatedMalware', [])
}
return enrichment_payload
```
**Implementation Challenges & Solutions:**
* **Rate Limiting:** Recorded Future APIs have rate limits. The middleware service implements a token bucket algorithm and caches frequently seen observables for a short period (10 minutes) to avoid redundant calls for repeated events from the same source.
* **Data Model Mapping:** Mapping Recorded Future's JSON response to our internal SIEM's Common Information Model (CIM) was necessary for consistent searching. We created a dedicated Splunk lookup table and a custom data model alias for the enriched fields.
* **Linkage:** The correlation ID is crucial. Every enriched event stored in Splunk contains the `correlation_id` of the original alert, allowing us to create unified views and dashboards.
The outcome is that analysts now see a single event in their investigation dashboard that combines our internal detection with external threat context. This has reduced the time to triage these specific alerts by approximately 70%, as the first step of manual lookup is eliminated. The enriched data also allows for more precise automation, such as escalating only alerts where a high Recorded Future risk score is accompanied by evidence of association with "Targeted Malware."
null