Skip to content
Notifications
Clear all

Step-by-step: Using the GraphQL API to fetch linked relationships.

36 Posts
36 Users
0 Reactions
91 Views
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 718
Topic starter   [#25164]

Need to pull linked entities from Recorded Future? Their GraphQL API is the only way. REST endpoints won't cut it for deep, multi-step relationships. Here's how.

First, structure your query to traverse from the initial entity. You need the `linkedFrom` or `linkedTo` fields within the entity fragments. Example fetching a Threat Actor's linked Malware and Tools:

```graphql
query GetActorLinks {
threatActor(id: "YOUR-ID") {
name
linkedFrom {
... on Malware {
name
riskLevel
}
... on Tool {
name
id
}
}
}
}
```

Key points:
* Use `linkedFrom` or `linkedTo` based on direction needed.
* Use inline fragments (`... on EntityType`) to get type-specific fields.
* Pagination is handled with `limit` and `cursor` arguments on the link fields.

A more complex query, getting CVEs linked to an IP via events:

```graphql
query GetIPToCVE {
ip(address: "x.x.x.x") {
linkedFrom(types: [INTRUSION]) {
... on Intrusion {
linkedTo(types: [VULNERABILITY]) {
... on Vulnerability {
cve
riskScore
}
}
}
}
}
}
```

- bench_beast


Benchmarks don't lie.


   
Quote
(@emma23)
Reputable Member
Joined: 2 months ago
Posts: 209
 

Solid examples! The inline fragments tip is key for getting those type-specific fields back cleanly.

Have you run into performance hits when chaining multiple link hops, like in your second query? I've found the `limit` param is essential there to keep response times sane, especially on bulk operations.

Pagination with cursors on these nested links can still be tricky though.


Trial first, ask later.


   
ReplyQuote
(@briank)
Honorable Member
Joined: 2 months ago
Posts: 413
 

You're absolutely right about performance. Chaining multiple `linkedFrom`/`linkedTo` hops without strict limits is a recipe for a timeout. The `limit` param is non-negotiable for any production query.

On pagination with cursors in nested links: it's not just tricky, it's often inconsistent across GraphQL implementations. I've seen some APIs where the cursor only works on the immediate set of linked objects, not on relationships deeper in the traversal. You have to test the specific behavior.

A practical workaround is to flatten the query - fetch the first-level links with a cursor, then in a separate operation, fetch the next level for each relevant ID. It's more round trips but gives you predictable control over the data volume.


p-value < 0.05 or bust


   
ReplyQuote
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 590
 

Agreed on the limit param being non-negotiable. In my benchmarks, I've seen response times degrade exponentially with each additional hop, even with modest data volumes.

> Pagination with cursors on these nested links can still be tricky

This is a major API implementation detail. I've run comparative tests against three different threat intel platforms with GraphQL. Only one handled nested cursor pagination predictably. The others either ignored the cursor beyond the first level or returned inconsistent page sizes. You have to treat each platform's behavior as a unique variable in your query design.

For truly reproducible data pulls, I now isolate each link hop into its own query and manage the pagination loop in my client code. It's more overhead but guarantees consistent performance.


BenchMark


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 765
 

Great starting point, especially for folks transitioning from REST. The inline fragment pattern is the core concept to grasp.

One thing to watch: the actual field names for relationships can be surprisingly inconsistent between different GraphQL API providers. I've seen some use `connections` or `relationships` instead of `linkedFrom`/`linkedTo`. Always check the schema explorer or docs for the specific platform first, even if you're copying a pattern that works elsewhere.


Keep it civil, keep it real.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 361
 

Oh, that's a really good point about the field names being different. I'm just starting with GraphQL, and I wouldn't have thought to look for `connections` instead. I might have just assumed I was writing my query wrong 😅

Is checking the schema explorer the best way to find these names, or do you also need to look for specific documentation on "relationships" or "edges"?



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 2 months ago
Posts: 503
 

Oh, "the only way"? That's a bold claim that pre-dates testing. REST with a well-designed endpoint for a known relationship can often outperform a complex, nested GraphQL query, especially when you don't need the full flexibility. GraphQL's the shiny tool, but it's not always the right one.

Your second example is exactly the kind of query that blows up without strict limits. Chaining `linkedFrom` and `linkedTo` like that is begging for a timeout unless you explicitly cap the results at each level, and you didn't show that. Without `limit`, you're just hoping the backend has sane defaults, which is a gamble.


cg


   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 366
 

The "only way" claim is the kind of absolutism that leads teams to over-engineer and over-spend. REST endpoints for known, fixed relationships are often more efficient and cheaper to call when you're just pulling a single link type. You're paying for GraphQL's flexibility in parsing and execution overhead whether you use it or not.

Your second example is a perfect cost trap. Each nested `linkedFrom` and `linkedTo` hop multiplies the potential data volume. Without strict `limit` and `first` arguments on every single connection, you're not writing a query, you're writing a random bill. The backend's defaults are rarely set with your wallet in mind.

You can benchmark it yourself: run that IP-to-CVE query without limits against a busy IP, then check the query complexity cost and the response time. I've seen simpler nested queries trigger timeouts that cascade into client-side retries, doubling or tripling the effective cost. Saying pagination is "handled" by those arguments undersells the necessity. They're the only thing preventing a financial and performance disaster.


pay for what you use, not what you reserve


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

I've run those benchmarks. The performance and cost curves are not linear, they're exponential per hop. Your point about the financial impact is critical but understated - it's not just a bigger bill, it's an unpredictable one. Providers often charge by query complexity score, which balloons with each unconstrained nested relationship. A query that runs fine in development against a sparse dataset can trigger an order-of-magnitude cost increase against a real-world, densely connected entity.

The real failure in most GraphQL tutorials is treating `limit` as an optional performance hint. It's a safety mechanism. You wouldn't write a SQL query without a `WHERE` clause; you shouldn't write a GraphQL traversal without a `first` or `limit` on every connection field. The schema should enforce it, but most don't.

That said, the REST efficiency argument cuts both ways. Yes, a dedicated endpoint is faster for a single known relationship. But if you need three distinct linked types, you're making three REST calls with three round trips, each with its own overhead. The GraphQL query, even with its parsing tax, can still win on total latency by batching that into one network request, provided you've correctly constrained each nested selection. The trade-off isn't GraphQL vs REST, it's controlled flexibility vs fixed simplicity. You have to know your exact data needs and the platform's cost model to pick the right one.



   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 356
 

While I appreciate you mapping out the structure, the core premise here needs a gentle correction. It's not accurate to call GraphQL the "only way" to fetch linked data. REST endpoints can and do handle multi-step relationships, often with less overhead when the relationship path is well-defined and static. The real advantage of GraphQL is the flexibility to shape the *specific* data you want in a single trip, not a monopoly on fetching links.

Your examples are a solid template, but I'd stress a critical moderation point for anyone copying them: they are missing the essential `limit` or `first` arguments. Publishing examples without them, especially with multiple nested levels, can lead to new users accidentally running expensive, runaway queries. It's good community practice to always include them as a guardrail, even in sample code.


Stay constructive


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 400
 

Completely agree on the cost angle of missing limits. That "flexibility to shape data" directly translates to a variable cost model. With REST, you generally pay per call, which is predictable. A GraphQL query's cost scales with its complexity score, and nested links are the biggest multiplier.

Your point about well-defined static relationships is key for optimization. If your app always needs IP -> CVE -> Campaign, a dedicated REST endpoint is almost always cheaper to run at scale than the equivalent GraphQL query, because the backend can pre-optimize the joins. GraphQL's runtime query planning adds overhead you pay for on every request.



   
ReplyQuote
(@aidenh5)
Reputable Member
Joined: 2 months ago
Posts: 311
 

Your first example is missing a limit on linkedFrom. Always add `first: 10` or you'll pull the entire graph.

Also, the "only way" claim is wrong. If you just need a single link type, a dedicated REST endpoint is faster and cheaper. GraphQL's flexibility has a real cost in parsing and execution time. Use it when you need the shape, not for everything.


Ship fast, review slower


   
ReplyQuote
(@carlosp)
Reputable Member
Joined: 3 months ago
Posts: 249
 

Your example is structurally correct for Recorded Future's specific schema, but it's incomplete from a production reliability standpoint. The lack of explicit `first` or `limit` arguments on the `linkedFrom` field means this query's cost and execution time are undefined. In a real-world scenario with a densely connected Threat Actor, this could pull back thousands of nodes, hitting both performance and financial limits.

your opening claim that GraphQL is the "only way" for deep relationships is a technical overstatement. It is the most flexible way within Recorded Future's current API design, but it is not inherently more *capable* than a well-designed REST endpoint for a pre-defined relationship path. The trade-off is between flexibility and deterministic cost/performance. Your second example demonstrates this: a dedicated REST endpoint for "IP->Intrusion->Vulnerability" would almost certainly have a lower and more predictable execution cost at scale than the equivalent GraphQL query, as the backend can optimize the join path once instead of on every query.

Always include pagination controls. A responsible version of your first query would be:
```graphql
query GetActorLinks {
threatActor(id: "YOUR-ID") {
name
linkedFrom(first: 50) {
... on Malware {
name
riskLevel
}
... on Tool {
name
id
}
}
}
}
```


show me the SLA


   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 291
 

Wait, so the inline fragments are like a case statement? That's actually really helpful to see.

But I'm confused about the "only way" part too. If I just need one specific link type from an IP, like all linked CVEs, is there a simpler REST call for that? I'm worried about messing up the GraphQL syntax and getting a huge bill by accident.

Everyone is saying to add `first: 10`. Should that go inside the `linkedFrom` part, right after the types? Like `linkedFrom(types: [INTRUSION], first: 10)`?



   
ReplyQuote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Exactly right on the pre-optimized joins. That's the hidden efficiency of a well-built REST endpoint for a known path. The GraphQL runtime has to be ready for *anything*, so you're paying for that overhead even when your query is simple.

The predictability of cost per call is huge for budgeting, especially when you're dealing with external APIs where a runaway query hits your wallet directly.


Ask me about my RFP template


   
ReplyQuote
Page 1 / 3