I'm trying to automate some threat intelligence pulls from Recorded Future into our own dashboard. The web interface is fine, but I need to pull entities and see how they're connected—like an IP address linked to a malware family, and then that malware linked to vulnerabilities.
The documentation mentions using the GraphQL API for "linked relationships," but the examples are pretty high-level. Can someone walk through the actual steps?
Specifically:
- What does the GraphQL query structure look like to go one level deep?
- How do you handle pagination when a node (like a threat actor) has dozens of linked entities?
- Is there a practical limit to how many relationship hops you can fetch in one query?
I'm used to REST APIs, so GraphQL is a bit new to me. A real-world example would help a lot.
Moving from REST to GraphQL for this is a smart move, especially for linked data. The key is thinking in terms of the graph structure, not separate endpoints.
For a one-level deep query, you'd start with the main entity and include a "links" or "connections" field in your selection. You'll need to know the specific link type name from their schema, something like `linkedMalwareFamilies`. The query asks for the fields you want on both the parent and the linked nodes in one nested structure.
On pagination, the API almost certainly uses cursor-based pagination for those connection lists. Look for `edges` and `pageInfo` objects in the response; you'll use the `first` argument to limit nodes per page and `after` with a cursor to get the next set. The practical limit on hops is usually about 4-5 levels deep before you risk a timeout or a massive response, but you can often get 2-3 hops comfortably by nesting those connection fields.
I'd suggest your first query just fetches an IP and its directly linked malware, keeping it simple. Once you see that structure, adding another level for vulnerabilities linked to that malware becomes clearer.
—daniel
The shift to GraphQL makes a lot of sense for linked data like this, but I get how the jump from REST can be tricky.
For a concrete one-level example, you'd structure your query to ask for both the IP entity and its connections in one go. Something like `ipAddress(/*args*/) { id, riskScore, linkedMalwareFamilies(first: 10) { edges { node { name, id } } } }`. You have to know the exact field name for the link type from their schema explorer.
On your other points, pagination is indeed cursor-based via `edges` and `pageInfo`. Start with a sensible `first:` limit. The hop limit depends on the API's complexity cost settings; hitting a depth of 4 or 5 is often where you'll start seeing timeouts or errors unless you're very selective with the fields you request at each level.
Raise the signal, lower the noise.