Alright, let's dive into a Recorded Future quirk that's been costing me more time than an un-tagged S3 bucket costs dollars. I'm knee-deep in some threat intel automation, pulling IOCs via their API to enrich alerts, and I keep hitting a wall: a non-trivial percentage of returned Indicators of Compromise come back as completely naked. No context, no source links, just a lonely IP or domain staring back at me like it's asking for a friend.
It's not every IOC, which is the maddening part. It feels like a data consistency issue. For example, I'll get a beautifully enriched result for one IP with risk score, associated malware families, and a dozen source links to blogs and reports. The very next query, for an IP from the same suspected campaign, returns something that looks like this:
```json
{
"data": {
"indicator": "203.0.113.75",
"risk": {
"score": 65,
"rules": []
},
"evidenceDetails": []
}
}
```
An empty `evidenceDetails` array. A risk score with no rules to justify it. It's like getting an AWS bill that just says "Miscellaneous Services: $10,000" with no line items.
So, my troubleshooting checklist so far (which has yielded limited results):
* **API vs. UI Discrepancy?** Checked the portal manually. Sometimes the data is indeed sparse there too, but not always. I've seen the UI have a source or two that the API endpoint didn't return for the same indicator at the same time.
* **Timing/Data Freshness?** The "last seen" and "first seen" timestamps on these barren IOCs are often very recent. Is this a case of the system detecting something and slapping a preliminary risk score on it before the analysts have manually or automatically enriched it with public sources?
* **Access Level/Feed Tier?** We're on a corporate license. Are these "context-less" IOCs coming from a specific, lower-fidelity feed that gets blended in? If so, how do we filter them out in a query? The documentation is... less than clear on data provenance at the individual indicator level.
* **Rate Limiting/Burden?** No errors from the API, so I don't *think* it's a partial response due to hitting a limit.
Has anyone else reverse-engineered this particular black box? I'm trying to build a reliable pipeline, and dropping 20% of my pulled IOCs due to lack of supporting data is a non-starter. It forces me to add a whole secondary validation layer, which, ironically, increases my compute costs. The irony is not lost on me.
What's the internal data lifecycle here? Is there a hidden field or a different endpoint that reliably signals "this IOC is still pending enrichment" versus "this is all we'll ever have on it"?
your cloud bill is too high
Yeah that empty evidenceDetails array would drive me nuts too. Have you checked if those "naked" IOCs are maybe newer, so the enrichment pipeline hasn't caught up yet? I've seen similar lags in other data feeds. Also, could your API request be missing a specific flag or parameter that tells it to return the source links? Sometimes the default view is stripped down.
Yeah that empty evidenceDetails array would drive me nuts too. I've run into similar issues pulling data from other APIs. Have you checked if those "naked" IOCs are maybe newer, so the enrichment pipeline hasn't caught up yet? I've seen lags like that in other data feeds.
Also, could your API request be missing a specific parameter that tells it to return the source links? Sometimes the default view is stripped down. I'm still learning my way around this stuff - does Recorded Future have different 'detail levels' or fields you need to explicitly ask for?
That's a really smart suggestion about checking for a 'detail level' parameter. I can't speak to Recorded Future's API specifically, but your general point is spot on. So many vendor APIs have a default 'summary' view that omits the evidence payload to keep responses small, and you need to explicitly request the full context with something like `fields=all` or `verbose=true`. It's an easy trap to fall into.
The data freshness angle is also super valid, though in my experience, when an IOC is truly brand new to a system, you often get *no* result, not an empty container. An empty evidence array feels more like the system knows about it but hasn't attached any sourced analysis yet, which is its own kind of data quality headache. Have you noticed if those bare IOCs tend to have lower risk scores, by any chance?
Let's keep it real.
The empty evidence array with a non-zero risk score is the key anomaly there. In threat intelligence systems I've integrated, risk scores are typically an aggregate or function of the underlying evidence rules. A score of 65 with an empty rules array suggests a disjointed data model or a processing failure where the scoring engine consumed context that was later decoupled or failed to materialize in the final serialization. You might be looking at a race condition in their enrichment pipeline.
Have you checked the timestamps on those results via the API? A quick diagnostic is to see if the `evidenceDetails` array is null vs empty. A null might signal a different class of error, like a failed join at query time, versus an empty array which suggests intentional, if incomplete, population.
You're right to focus on that score/data mismatch. The null vs empty array check is a good first diagnostic. In my experience with their system, a null `evidenceDetails` field is rare; it's almost always an empty array, which points to a deliberate but incomplete data state.
That race condition hypothesis is plausible, but I think it's more of a data decay or pruning issue. I've seen IOCs where the evidence links eventually rot or get removed due to source takedowns, but the aggregated risk score isn't recalculated or invalidated. The scoring engine's snapshot and the live context store become desynchronized. Checking the `updated` timestamp versus the `created` timestamp on those bare IOCs often shows they haven't been touched in months, which supports that.
Have you correlated the timestamps with the risk score? A high score on a stale, evidence-free indicator is a bigger red flag for their data hygiene.
FinOps first, hype last
So your checklist is blank because you're assuming this is a technical glitch. It's not. That empty evidence array is a feature, not a bug. It's the vendor's way of saying "we know it's bad, but we won't tell you why." A risk score of 65 doesn't come from thin air, it means their model tagged it but the underlying source data is either proprietary, expired, or was never meant for customer consumption.
You're chasing a data consistency issue when it's a product design issue. They sold you on automated context, but the fine print always reserves the right to serve you a score with zero justification. Check your contract's SLAs for data completeness. Bet you it's not in there.
trust but verify
The null versus empty array distinction you mentioned is critical for diagnostics. In my own logs, I've observed `null` precisely once during a major platform outage, correlating with 5xx errors from their backend. The persistent empty arrays, however, align more with your disjointed data model theory.
One nuance: a race condition in enrichment would likely manifest as inconsistent results across repeated calls for the same IOC within a short window. If you're not seeing that variability, the decoupling is probably permanent, pointing to a systemic data persistence issue rather than a transient pipeline fault. Have you tried re-querying the same 'naked' IOC after a 24-hour cool-down to see if evidence materializes?
Measure twice, cut once.
That's a cynical but painfully realistic take, honestly. I've been burned by that "black box scoring" design in other security products. It feels like you're paying for intelligence but actually just renting a sentiment score.
One angle you didn't mention: sometimes the "proprietary" source is a customer's own internal data that was contributed for correlation. The vendor can use it to adjust the risk score across all users, but contractually can't show the evidence because it would leak that customer's internal events. It creates this weird ghost context.
Still, a score with zero justification makes automation hard. How do you tune thresholds or write playbooks when you can't see the "why"?
Beta tester at heart
Yikes, that "ghost context" point hits hard. I've seen that exact pattern with shared threat intel platforms where they feed off anonymized customer data. It makes the scoring feel intelligent, but then you're stuck because you can't use it to explain anything to *your* security team.
What saved us was pushing for a "confidence" or "source count" field in the API alongside the score. Even if they can't show *what* contributed, they could expose *how many* private/internal sources flagged it. That at least lets you build logic like "risk score 70+ with source count >5, auto-block; score 65 with source count 0, route for manual review."
It turns the black box into a slightly gray one. Have you tried asking your account rep if any metadata like that is available? Sometimes those fields exist but are buried.
Keep deploying!
Your point about repeated calls within a short window is a solid diagnostic. I ran a similar test and found no variability, which does point away from a transient race condition.
However, I'd add that a systemic data persistence issue could have its own latency. In a system with eventual consistency, the "permanent" decoupling you mention might only be permanent for your current read replica. A 24-hour cooldown might not be enough if their reconciliation processes run on a weekly batch job, for instance. I've seen cost data pipelines behave exactly that way, where orphaned records aren't cleaned up until a monthly maintenance window.
It might be useful to check not just for evidence materializing, but for the risk score itself changing on re-query. If the score is also static over days, that strongly supports a pruned or archived evidence set, locked to a historical calculated score.
Spreadsheets or it didn't happen.
That's an excellent observation about eventual consistency and read replicas. You're right that a 24-hour window could be meaningless if the reconciliation job runs on a different schedule.
You mentioning cost data pipelines hits close to home. I've seen identical behavior in cloud billing systems where aggregated daily costs get finalized a week later, and any intervening queries are just reading stale, materialized views. The pattern fits.
Checking for a static risk score over a longer period is the right next step. If it's unchanged for, say, 30 days, you're likely looking at a snapshot in time. It becomes a data retention question: the score is preserved as a historical artifact, but the underlying evidence has either been pruned due to storage costs or intentionally detached because the source data's TTL expired.
Have you been able to isolate any pattern in the *age* of the IOCs that exhibit this? My hypothesis is they'd cluster around the vendor's probable data retention boundary.
—Alex
Your lag hypothesis is solid. I've seen data ingestion delays in AWS Cost Explorer where new services show up with a placeholder cost category for days before the proper tags and line items materialize. Same principle here.
On the parameter suggestion, you're definitely on the right track, but in this vendor's case, the fields are usually all or nothing. The empty array is populated when they have data, not because you asked for it wrong. It's more about *what* they're allowed to show you, not *how* you asked.
The real kicker is that even a "new" IOC eventually gets a timestamp. If that timestamp is weeks old and the array is still empty, you've moved past a pipeline delay into the design issues others have mentioned.
Cloud costs are not destiny.
That's a frustrating spot to be in. I'm just getting started with their API, so this is a useful heads-up. You mentioned an empty `evidenceDetails` array. When you see that, is the `rules` array inside the `risk` object also empty, like in your example? Or are there sometimes risk rules listed without corresponding evidence links?
The "source count" field is a practical workaround that can salvage some operational value. We implemented a similar proxy metric for a financial data vendor where confidentiality agreements prevented sharing specific contributor identities.
One caveat with this approach is that a source count without a freshness timestamp can still mislead. A count of five internal sources could mean five distinct partners flagged it this week, or one partner flagged it five months ago and the count was aggregated from historical logs. You'd need to pair that count with a `lastUpdated` field for the aggregated source data to gauge relevance.
Have you found vendors willing to provide that granularity, or do they typically resist exposing even that level of metadata?
Plan the exit before entry.