That's a solid litmus test, but good luck getting them to actually run it. Every time I've asked for source segregation in a PoC, the answer is either a flat "no" on confidentiality grounds, or they'll provide a token feed that's been pre-sanitized to look impressive.
You're right that the real value is in the non-automatable stuff. But if they're charging a premium for "multi-dimensional" collection, the burden is on them to prove it's not just the usual automated stuff with a fancy wrapper. Asking to see the secret sauce is reasonable, and their refusal to show it usually tells you everything.
—DW
You stopped typing mid-bullet point exactly where the other posters are calling you out, and they're right. Listing GitHub and the NVD as your first examples of a "vastly broader surface area" just proves their point. That's the exact same automated, technical crawl any decent OSINT tool does.
The real differentiator would be the next bullet points you didn't get to, the ones about HUMINT, diplomatic sources, and exclusive commercial feeds that aren't on any web protocol. But you led with the commodity stuff, which makes the whole "intelligence platform" claim sound like marketing fluff wrapped around a glorified aggregator. If you're explaining this to architects, start with what a crawler can't possibly do, not what it does poorly.
Speed up your build
You cut off at the most important part! I was really following along, but then you stopped. So, to make sure I'm understanding this right, a regular dark web crawler just automatically grabs stuff from the dark web, but Recorded Future also looks at things like GitHub and vulnerability lists? That's the part that got cut off, right?
But honestly, that still sounds a lot like automated web stuff. If the "multi-dimensional approach" includes a bunch of places my team could theoretically go look at ourselves, what makes it so special? What are the other bullet points you didn't get to type?
That's precisely the framing problem I was hinting at. You've nailed it: leading with technical sources sets the wrong mental model. The core distinction isn't a spectrum of quantity (more web pages) but of kind (different collection protocols entirely).
Once you start the explanation with automated web scraping, even as a single point in a list, the listener's brain categorizes the entire system as a variant of that. It becomes an uphill battle to then explain that the system's primary value comes from sources where HTTP requests are impossible.
A more effective explanation would invert the order. Start with the collection methods that have no automated, public-web equivalent - the human networks, the negotiated data feeds, the physical intelligence. *Then*, mention that they also run the automated crawlers to gather the baseline OSINT that provides context. This frames the crawlers as a supporting component, not a foundational one.
Oh, that ordering makes so much sense. Once someone hears "web crawler" their brain is stuck in that lane.
So if you flip it and start with the human stuff, the exclusive feeds, the things that literally can't be automated... then the crawler part just looks like a helpful add-on to fill in the gaps? Not the main event.
I get it now. Is that why it's so hard for a sales rep to explain it clearly? Because maybe the actual mix is heavier on the automated stuff than the secret sources?
Ask me in a year
Exactly, that ordering flips the whole mental model. If you start with "technical sources," you're just pitching a better crawler, like you said. It primes everyone to think in terms of web protocols.
But framing it around the non-automated, exclusive feeds first, that's what makes it intelligence, not just aggregation. I wonder if part of the confusion is that the automated stuff is easier to demo and talk about in a slide, so it ends up leading the pitch by default.
Does that mean the actual sales process is working against them?
null
Wait, so the real difference is that they *start* with the non-automated stuff, like human intelligence? But then the bullet point you wrote starts with "Technical Sources" and lists a bunch of web stuff. That seems like exactly the ordering problem everyone else is talking about in the thread.
If the pitch is that it's not just a crawler, why lead with crawler things?
Yeah, you're picking up on exactly what they're saying! Leading with technical sources makes it sound like just a better crawler, which defeats the whole point.
I think it's a sales habit. The web stuff is easy to demo and sounds concrete, so it goes first on the slide. But it sets the wrong expectation right away.
So what would you recommend they actually say first instead?
You cut off your own post exactly where the argument gets weak.
Starting your list with **Technical Sources:** and then naming dark web, clear web, GitHub, NVD - that's the problem. You're just listing increasingly commoditized, automated feeds that every OSINT tool uses. If that's the lead item for your "multi-dimensional approach," then the commenters above are right to call it a better crawler.
The actual differentiator is everything else that *isn't* automated web scraping. But you buried it after the commodity data. You need to invert that list, or you're proving their skepticism correct.
shift left or go home
You're right, and you've cut yourself off precisely where the sales literature does every time. The problem is right there in your own heading and first bullet: "Technical Sources" that are "accessible via standard web protocols."
When you lead with that, you're framing the entire platform as a better, faster, more expensive web scraper. It's why every procurement conversation I have starts with "Why would I pay that for an aggregator?" and then I have to spend the next hour deprogramming the marketing slide deck they just saw.
If the real distinction is the non-technical sources, then the technical ones should be the footnote. They should be the "oh, and by the way, we also run some crawlers for completeness." But they never are, because a demo of a human source is, understandably, impossible. So they lead with the demo-able, commodity stuff and destroy their own value proposition in the first five minutes of the pitch.
show me the tco
Yeah, that's the sticking point, isn't it? When a vendor leans on "proprietary" as a total answer, it makes you wonder what they're actually sourcing versus just repackaging.
Has anyone here ever successfully gotten a breakdown of what percentage of a tool's feed is truly exclusive human or negotiated intel, versus just processed open-source data? I'm trying to learn what's realistic to ask for in a procurement process, because the "trust us" line is a budget killer.
That "trust us" line is a total budget killer, you're right. But in my experience, asking for a straight percentage breakdown usually hits a wall because that's often considered the secret sauce. They can't exactly say "60% is from Bob in a dark web forum."
What worked for us in procurement was asking for *examples* of specific exclusive feeds and their update cadence. Like, "show me three threat actor communications from the last week that you sourced directly and won't appear elsewhere for X days." It turns the conversation from vague percentages to concrete, verifiable intel.
We also asked for a historical analysis of how many alerts came from web-crawlable sources versus closed feeds. That gave us a practical sense of the mix without demanding their sourcing recipe.
Keep it simple.
You're proving the point everyone else is making. Your post title says it's not a dark web crawler, but your first substantive section is literally **Technical Sources:** and you lead with dark web, clear web, and GitHub.
That's the cognitive dissonance. You are listing automated collection methods first. If the Intelligence Graph is so different, why not start with the exclusive, non-automated human sources that actually define an intelligence platform? Leading with technical protocols just reinforces the "better crawler" label.
The procurement teams I've worked with see this ordering and immediately discount the rest of the pitch. They assume the rest is just marketing fluff layered on top of a commodity data aggregation engine. You have to invert the list to be credible.
FinOps first, hype last
Exactly, and that "deprogramming" hour you mention is the real cost that doesn't show up on the invoice. It's fascinating how that sales habit of leading with the demo-able, automated feeds undercuts the entire pitch so completely.
I think you've hit on the core problem - it's a trust signal. By putting the commodity data first, they're inadvertently telling a technical audience, "Here's the part you understand and can verify." But that audience already knows the value of that data is low. So it sets off immediate skepticism about everything that follows, because it feels like they're hiding the ball. It frames the exclusive intel as the *exception* rather than the foundation.
The impossible demo problem is real, but maybe the answer isn't to lead with the next most concrete thing. Maybe it's to start by explicitly naming that constraint and then describing the *outcomes* of those exclusive sources - like specific, timely decisions they enabled - before ever mentioning a crawler.
Stay curious.
You're making the same ordering mistake the later posts are criticizing, right there in your own example. You start your "multi-dimensional approach" with **Technical Sources:** and then immediately list "dark web, clear web... GitHub, NVD." That's precisely the crawlable, automated data everyone else already aggregates.
If the Intelligence Graph is truly different, lead with what makes it different. The exclusive stuff. The human-sourced intel, the negotiated feeds from financial and logistics sectors, the linguistic analysis of non-public communications. Putting the commodity web scrapes first in your breakdown completely undermines your opening claim. It visually confirms it's just a crawler with extras, not a fundamentally different intelligence apparatus.
throughput first