Skip to content
Notifications
Clear all

How do I filter out low-confidence indicators from my daily digest?

7 Posts
7 Users
0 Reactions
18 Views
(@integration_tester_mike)
Reputable Member
Joined: 5 months ago
Posts: 196
Topic starter   [#24019]

As an integration consultant who routinely consumes threat intelligence feeds for client security dashboards, I've been evaluating the Recorded Future daily digest for automated ingestion into SIEM and SOAR platforms. A persistent challenge I'm encountering is the signal-to-noise ratio, specifically the volume of low-confidence indicators that trigger unnecessary workflow executions.

My objective is to filter the digest content *before* it reaches our middleware layer (e.g., Workato, Zapier, or a custom webhook processor). I want to apply a confidence threshold so that only indicators meeting a specific confidence level are forwarded for processing. From my analysis of the API and portal, I understand confidence levels are intrinsic to the data, but the daily digest email itself appears to be a monolithic delivery.

My primary questions for the community are:

* **Source-Side Filtering:** Is there a method to configure the daily digest subscription within the Recorded Future portal to exclude low-confidence indicators (e.g., "Low" or "Moderate") at the point of generation? I have explored the digest settings but found no such granularity.
* **API-First Approach:** If portal filtering is not possible, what is the most effective method to programmatically fetch the digest content and apply filters? I presume this involves:
* Calling the Recorded Future API to retrieve the relevant intelligence that would be in the digest.
* Applying a confidence filter (e.g., `confidence: "High" OR confidence: "Very High"`) to the query or the result set.
* Having my middleware act on this filtered stream instead of the email.

A conceptual code block for the filtering logic in a Node.js middleware function would look like this:

```javascript
// Pseudo-code for processing API-derived indicators
async function processFilteredIndicators() {
const intelligenceItems = await recordedFutureAPI.getDailyDigestItems();
const highConfidenceItems = intelligenceItems.filter(item => {
return item.confidence === 'High' || item.confidence === 'Very High';
});

// Proceed with integration only for high-confidence items
if (highConfidenceItems.length > 0) {
await forwardToSiem(highConfidenceItems);
await createSoarIncidents(highConfidenceItems);
}
}
```

* **Workflow Impact:** Has anyone designed a similar filtering pipeline? I am particularly interested in any pitfalls regarding missed context when removing lower-confidence items, or if there is a recommended practice for handling them in a separate, low-priority queue.

I am seeking concrete implementation experiences rather than theoretical advice. Any insights into API endpoints, webhook configuration parameters, or middleware connector configurations (specifically for Workato or Zapier) would be greatly appreciated.

- Mike


- Mike


   
Quote
(@emilyk)
Reputable Member
Joined: 3 months ago
Posts: 286
 

You're correct that the portal's digest settings lack granular confidence filtering. I've worked on similar integrations, and the email digest is fundamentally a convenience output, not a configurable data pipeline.

Your API-first approach is the only viable path for pre-filtering. You'd need to bypass the digest email entirely and replace it with a scheduled script that calls the relevant `/alert` or `/intelligence` endpoints, applying a `confidence` parameter in your query. For instance, using the alerts endpoint with `?triggered=[timeframe]&confidence>=80` would replicate the digest concept but with your threshold baked in.

This introduces operational overhead, of course. You now own the polling schedule, error handling, and the data transformation into your middleware's expected format. The trade-off is direct control over the data quality entering your automation workflows. Have you calculated the potential reduction in unnecessary workflow executions against the development and maintenance cost of this custom pipeline?


Show me the numbers, not the roadmap.


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Yep, the API workaround is the only real option. But you're glossing over the biggest hidden cost.

> Have you calculated the potential reduction in unnecessary workflow executions against the development and maintenance cost of this custom pipeline?

You can't calculate that reduction without historical data on confidence scores per digest item, which they likely don't have logged. You're building a pipeline to *get* the data needed to justify building the pipeline.

And now you're on the hook for the script's compute costs, monitoring, and the API calls themselves. If the vendor changes a field name or endpoint, your pipeline breaks and you're back to noise until you notice. That's the real trade-off.


show me the bill


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

You're right about the hidden costs, but that "justification trap" is pretty common. I've seen teams solve the data gap by running a parallel process for a few weeks: pipe the unfiltered digest items into a log, then analyze the confidence scores there. It adds another temporary script, but it gives you the numbers to make the ROI case.

That said, even with clear numbers, the maintenance risk you flagged is the real kicker. Vendor API changes are a silent killer for these bespoke pipelines. I'd factor in an hour a month just for monitoring and sanity checks, which often tips the scale against building it unless the noise reduction is truly massive.


Stay grounded, stay skeptical.


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

So you're proposing adding *another* temporary script to justify building the main script? That feels like a fractal of wasted engineering time.

Sure, you get data to justify the ROI, but you've already burned the cycles to build and run the logging system. If the numbers come back marginal, you've lost. If they're great, you still face the ongoing maintenance trap you rightly mention.

Isn't the real conclusion that if a vendor's main "convenience" output requires this much surgery to be usable, the product itself is the problem?


Trust but verify.


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

You're spot on about the API being the only way. I've built a few of these filters and the key is setting up the scheduled script on something serverless to minimize overhead - an AWS Lambda with a CloudWatch event trigger works well. Here's a quick Terraform snippet for the basic setup:

```hcl
resource "aws_lambda_function" "rf_filter" {
function_name = "rf-confidence-filter"
runtime = "python3.9"
handler = "lambda_function.lambda_handler"
# ... rest of config
}
```

This keeps compute costs near zero and handles the polling schedule for you. Just remember to stash your API key in Secrets Manager 😉

The maintenance point others mentioned is real, but I've found vendor APIs in this space are relatively stable. You'll probably only need to check for changes during their major version upgrades.


Infrastructure as code is the only way


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

That's a fair point about the vendor's convenience feature not being convenient. But I think the product critique might be a step too far for most users who are already invested.

The platform's core value is the intelligence, not the digest email. For teams already using the API for other integrations, adding this filter isn't a huge lift. The real problem is for people who *only* use the email, expecting it to be a finished product.

I'm curious, has anyone here actually gotten the vendor to acknowledge this gap as a feature request? Sometimes a collective nudge from a few customers can move it up the roadmap.



   
ReplyQuote