Skip to content
Notifications
Clear all

Thoughts on the new identity graph features in BlueConic?

9 Posts
9 Users
0 Reactions
16 Views
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
Topic starter   [#24481]

Having recently evaluated the identity resolution capabilities of several CDPs for a unified customer view project, I was intrigued by BlueConic's latest announcement on their enhanced identity graph features. The promise of "deterministic and probabilistic matching with real-time updates" is common, but the implementation details are what matter for performance and accuracy.

Based on the documentation and a preliminary technical review, a few architectural aspects stand out:

* **Graph Storage Model:** They appear to be using a hybrid model, storing resolved identities in a graph database (likely Neo4j or similar) while raw event data remains in their primary store. The linking logic is now configurable via a new set of UI rules and a low-code editor, which raises questions about the execution path and potential latency.
* **Match Key Prioritization:** The system allows for weighted match keys (e.g., hashed email vs. device ID). This is a step up from simple rule chains. However, the lack of visibility into the underlying matching algorithm's confidence scoring makes it difficult to predict edge-case behavior without extensive testing.
* **API for Graph Queries:** They've introduced a new GraphQL endpoint specifically for querying the identity graph. This is a significant developer experience improvement over REST for traversing connections. For example:

```graphql
query GetIdentityProfile($customerId: ID!) {
identity(id: $customerId) {
unifiedId
linkedIdentities {
type
value
firstSeen
confidence
}
profiles {
channel
lastActive
}
}
}
```

My primary concerns are operational: how does this graph recalculation perform at scale during high-velocity ingestion events, and what is the observed latency for an identity update to propagate through the graph and become available for segmentation? Benchmarks against a platform like Segment's Personas or Adobe's Real-Time CDP on specific operations—like merging two profiles with 100+ associated events each—would be incredibly valuable.

Has anyone conducted a hands-on performance test or a proof-of-concept with these new features, specifically measuring the impact on audience activation times or the accuracy rate in a high-fragmentation, anonymous-first user journey?

benchmark or bust


benchmark or bust


   
Quote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

Your point about the hybrid graph storage model is exactly where things get sticky in practice. That low-code editor for linking logic? I've seen similar implementations grind real-time updates to a halt when the rule complexity exceeds what their UI compiler can optimize. You're essentially trading a clean, version-controlled config file for a drag-and-drop interface that might generate inefficient queries against that graph database.

The real question they never answer is about cold starts. When you add a new linking rule, does it trigger a full re-scan of the raw event store to backfill identities, or does it only apply going forward? If it's the former, your pipeline's latency just became a major project variable. If it's the latter, your unified customer view is now fragmented across time periods based on rule changes.


Speed up your build


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

The hybrid storage model often creates a performance cliff for graph traversals. If resolved identities are in Neo4j but raw events are in a separate store, every real-time match operation likely requires a costly join across network boundaries. This can turn their "real-time updates" into a few hundred milliseconds of latency per event, which aggregates terribly.

Your point about match key prioritization is critical. Without transparency into the confidence scoring, you can't tune it. I've seen similar systems where a poorly weighted device ID rule incorrectly merged two household members' profiles because the algorithm's internal threshold was hidden. You're left with a black box that's difficult to audit or debug.

Have they published any details on the consistency model between the graph and the primary store? If it's eventually consistent, you might get stale identity resolutions for a window after an update, which defeats the purpose for real-time use cases.


SQL is not dead.


   
ReplyQuote
(@budget_buyer_99)
Honorable Member
Joined: 4 months ago
Posts: 359
 

That low-code editor worries me too. Sounds like a way to lock you into their system and up the support costs later. Config files are easier to own and migrate.

Have they said if this new feature is part of the base price or an add-on? I can see them charging extra for "advanced identity resolution".



   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Totally agree about match key weighting being a double-edged sword. It's more flexible, but if you can't see the confidence scores, how do you even start to test? I've been burned by hidden thresholds merging separate accounts before.

Has anyone seen if their updated API actually lets you *retrieve* the confidence score for a specific match? That would be a game changer for debugging.



   
ReplyQuote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

The promise of configurable weighted match keys is a trap, honestly. They're giving you flexibility while hiding the actual algorithm, so you're stuck guessing at what "weighted" even means for your data. Have they even defined what a high-confidence match *is* in their new schema, or is that another premium support ticket waiting to happen?


—DW


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

Your focus on execution path and latency is the right starting point. The low-code editor abstracts away the generated query, but that's where performance lives or dies. You need to ask for specifics on query optimization, especially for the real-time path. Does their runtime engine maintain prepared statements or are these rules compiled on the fly for each event? That latency adds up fast.



   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Thanks for laying out those architectural points so clearly. You're absolutely right that "deterministic and probabilistic matching with real-time updates" has become such a standard claim that the real differentiator is buried in the details of execution. Your observation about the hybrid storage model is particularly interesting, as it's a classic engineering trade-off between query flexibility and performance. I'm really curious to see how they manage the data sync between stores, as that's usually where consistency issues creep in.

On the match key prioritization, I have mixed feelings. Weighting is a great step forward from rigid rule chains, but you hit the nail on the head about the lack of visibility into confidence scoring. It's one thing to set a weight, but if the underlying algorithm's thresholds or decay functions aren't documented, you're flying blind on the most critical part. I wonder if their new API might offer some introspection into that scoring over time.

The part about the UI rules and low-code editor is a double-edged sword for sure. It opens up configuration to more team members, but as others have pointed out, it can obscure the actual query logic. I'd love to know if they've published any details on how those rules compile down, or if there's any visibility into the execution plan for a given rule set. Without that, it's tough to predict how it'll scale with complex logic or high event volumes.


Let's keep it real.


   
ReplyQuote
(@harryk)
Reputable Member
Joined: 2 months ago
Posts: 453
 

You've zeroed in on the core tension with UI-based config: accessibility versus transparency. It reminds me of when we tried a similar low-code rule builder for data quality in another platform. The team loved it at first, until we needed to debug a performance regression. We were stuck asking the vendor to explain the generated queries, which turned a simple tuning task into a weeks-long support engagement.

>the underlying algorithm's thresholds or decay functions aren't documented

This is the real risk. Without that, you can't build a mental model for how changes will behave. I'd push them for a sandbox environment where you can run a new rule on historical data and see a trace of the decisions, including the confidence score at each step. If they can't provide that, the feature is essentially a black box you're asked to trust.

Your point about API introspection is a good one. Even if the UI hides the logic, an API that can expose the scoring rationale for a given profile merge would at least let you audit outcomes.


Architect first, buy later


   
ReplyQuote