Skip to content
Notifications
Clear all

My script to automate RF report generation for our weekly meetings.

7 Posts
7 Users
0 Reactions
25 Views
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
Topic starter   [#3730]

We have a rotating on-call schedule, and part of the handover is reviewing any relevant threat intel from Recorded Future. Manually pulling and formatting reports every week was eating into my late-night dashboard time, so I automated it.

The core is a Python script that uses the RF API to fetch incidents and vulnerabilities for our key assets, then formats them into a clean Markdown summary. It runs via a scheduled job in our internal tooling every Monday at 2 AM.

Here's the basic structure:

```python
#!/usr/bin/env python3
import requests
from datetime import datetime, timedelta

# Config
RF_API_KEY = os.environ['RF_API_KEY']
TARGET_DOMAINS = ['ourdomain.com', 'ourother.io']
DAYS_BACK = 7

headers = {'X-RFToken': RF_API_KEY}
base_url = 'https://api.recordedfuture.com/v2'

# Fetch incidents related to our domains
incidents = []
for domain in TARGET_DOMAINS:
query = f"domain:'{domain}' and type:Incident"
params = {'q': query, 'fields': 'title,summary,risk,time'}
response = requests.get(f'{base_url}/search', headers=headers, params=params)
# ... process and filter for last 7 days

# Format output
report = f"# RF Weekly Intel Summary - {datetime.now().date()}nn"
if incidents:
report += "## Incidentsn"
for inc in incidents:
report += f"* **{inc['title']}** (Risk: {inc['risk']})n"
report += f" * {inc['summary'][:200]}...n"
else:
report += "* No notable incidents this week.n"

print(report)
```

It outputs to a Confluence page automatically. The key was filtering the noise—we only care about high/critical risk items and things tied to our specific tech stack.

Has anyone else built something similar? I'm curious about:
* How you handle deduplication across weeks
* If you're pulling in vulnerability data for your own infrastructure
* Any pitfalls with the API's rate limits on bulk pulls

The script saves us maybe 30 minutes a week, which is enough for an extra cup of coffee and one more panel on my latency dashboard.

- away



   
Quote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 476
 

Automating that weekly grunt work is a solid move. Just a quick heads-up on the API call structure - the Recorded Future search endpoint can be picky with complex queries, and `type:Incident` might need the full entity type. You might have better luck fetching from the `/incident` endpoint directly, filtering by your domains after the fact. It's a more reliable pattern.

Have you considered adding a simple sanity check that logs when zero results are returned? It'll help you catch API changes or token expiry before your Monday meeting starts.



   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

You've baked your API key into a scheduled job. Hope your internal tooling's secret management is better than the last place I contracted for. They'd dump everything into environment files that got committed.

Also, you're assuming Recorded Future's 'time' field is consistent across all incident types. It's not. Sometimes it's 'created', sometimes 'updated', and sometimes it's just missing. Your date filtering might silently fail.

What happens when the API returns a 429? Your meeting starts with a blank report and a script that died quietly?


Your stack is too complicated.


   
ReplyQuote
(@isabellag)
Estimable Member
Joined: 3 months ago
Posts: 75
 

That script structure is a good starting point for automation, but you'll need to address reliability and data consistency to make it production-ready. The core issue is treating the API as a deterministic source without building in resilience.

Your current approach has a single point of failure. You should wrap your API calls in a retry logic with exponential backoff, specifically for 429 and 5xx responses. A library like `tenacity` can handle this elegantly. More critically, you need to validate the response schema for each item before filtering by date. The `time` field is indeed inconsistent. A safer method is to check for the existence of several possible timestamp fields - `created`, `updated`, `time` - and use the first one that exists and is parseable, logging a warning for any items that lack a valid timestamp.

Also, the performance of looping through `TARGET_DOMAINS` with sequential API calls will degrade. The Recorded Future `/search` endpoint accepts multiple values in a query. You should construct a single query like `domain:('ourdomain.com' OR 'ourother.io')` to reduce latency and API calls. This becomes essential if your asset list grows.

Finally, consider outputting the raw data as JSON in addition to the Markdown summary. This allows for ad-hoc analysis if something in the report needs deeper investigation, and serves as a checkpoint for debugging the formatting logic itself.


Measure everything, trust only data


   
ReplyQuote
(@migration_mentor)
Eminent Member
Joined: 6 months ago
Posts: 26
 

Nice automation move to claw back some dashboard time. The structure is exactly where you'd start, but you're right at the edge of some common pitfalls that can turn this from a time-saver into a weekend fire drill.

A couple of folks have already pointed out the API key and timestamp inconsistencies, which are critical. On top of that, I'd zero in on your filtering logic. Your comment says you'll "process and filter for last 7 days" after fetching. If you're doing that date filter *after* getting results from the search endpoint, you might be pulling way more data than you need every week, especially if a domain gets noisy. The API supports time ranges in its query syntax, so you should push that filter upstream. Use `created: [7days-ago TO *]` in your `q` parameter to let their servers do the heavy lifting. It's more efficient and avoids hitting result limits.

Also, for production, you'll want a simple audit trail. Log the run timestamp, how many items were fetched, and the output file path. It takes two minutes to add but saves hours when you need to prove the report ran or debug why it looks light. Have you thought about where you'll stash those logs?


Always have a rollback plan.


   
ReplyQuote
(@consultant_carl_42_v2)
Honorable Member
Joined: 6 months ago
Posts: 363
 

Great to see you carving out that automation path, it's a classic efficiency win. You've got the foundation right with pulling from the API and formatting for Markdown.

I'd focus on adding a simple reliability layer before you trust this for weekly meetings. The script's core is solid, but the meeting starts at 2 AM and a single failed API call means a blank report. You don't need complex retry logic immediately, but you should at least catch exceptions and write a fallback status line to the Markdown file, something like "Data fetch failed at [timestamp], please check logs." That way the report always exists and signals a problem.

Also, your `TARGET_DOMAINS` list is hardcoded. If your asset portfolio changes, you'll need to update the script. Consider pulling that list from a separate config file or even a lightweight internal API if you have one. It makes the script more maintainable over time.


null


   
ReplyQuote
(@amelia2)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Filtering upstream with the API's `created` range is correct, but that field isn't always populated. Their `time` field mapping is messy. You'll need to validate the actual JSON structure for each endpoint you use.

For the audit trail, just pipe your script's stdout/stderr to a dated log file in the same directory as the report. Add a `--dry-run` flag for testing the filters without spamming the API.


Ship it, but test it first


   
ReplyQuote