Skip to content
Notifications
Clear all

Help: my CDP identity graph is producing way too many merged profiles - why?

7 Posts
7 Users
0 Reactions
35 Views
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
Topic starter   [#21773]

Hey folks, hoping to get some collective wisdom here. I’m helping our marketing team troubleshoot our CDP's identity resolution, and we're seeing a huge number of merged customer profiles that just don't seem right. Our "identity graph" is lumping together what are clearly different people, which is throwing off our segment counts and activation.

We’re using a deterministic rule set based on email and user ID, with some fuzzy matching on IP addresses and device IDs for unknown visitors. The volume of merged profiles has spiked about 300% in the last month. 😬

Here’s a simplified version of the rule configuration we’re using:

```json
{
"identity_rules": [
{
"priority": 1,
"rule_type": "deterministic",
"attributes": ["email", "user_id"],
"action": "merge"
},
{
"priority": 2,
"rule_type": "fuzzy",
"attributes": ["ip_address", "device_id"],
"threshold": 0.8,
"action": "merge"
}
]
}
```

From my observability world, this feels like an overly sensitive correlation rule creating noise. I’ve been checking the raw events feeding the graph, and I'm wondering:

* Could shared or dynamic IPs (like from a corporate network or mobile carrier) be causing the fuzzy rule to over-match?
* Is there a common pitfall in how device IDs are being generated or persisted across different platforms (web vs. app)?
* Should we be adding more negative correlation rules to prevent merges when certain high-confidence fields (like last name) don't match?

I’d love to hear from anyone who has tuned these rules before. What metrics did you track to validate your merges? Did you have to dial back fuzzy matching significantly? Any dashboard screenshots of your merge rates over time would be super helpful!


Dashboards or it didn't happen.


   
Quote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Your fuzzy rule threshold at 0.8 is way too aggressive for IP and device ID. That's almost certainly the culprit.

Shared corporate or public IPs, plus mobile devices that share ad IDs, will hit that similarity score easily. It'll merge whole office buildings or families into single profiles.

Drop the fuzzy threshold to at least 0.95 and check your match keys. You likely have a data quality issue where placeholder or null device IDs are being scored as identical.


Prove it with a benchmark.


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

user1054's point about the fuzzy threshold is correct, but the 0.95 suggestion could still be problematic. The core issue is using fuzzy matching on inherently unstable identifiers for deterministic merging.

An IP address isn't a fuzzy identifier, it's a categorical one. Two users sharing a corporate NAT have identical IPs, not similar ones. Fuzzy logic on a full IP string gives you a false sense of precision. A better approach is to separate the logic: exact match on a hashed identifier, then apply a completely different rule for shared-network scenarios based on a confidence weight, not a similarity score. Device IDs should be handled the same way - either they match exactly or they're different.

This often points to a rule ordering problem. If your deterministic email/user_id rule runs first and creates a profile, then your fuzzy IP rule runs and merges another profile onto it, you've now contaminated your graph. You need to quarantine matches from low-fidelity signals, not merge them directly.


brianh


   
ReplyQuote
(@emma23)
Reputable Member
Joined: 3 months ago
Posts: 212
 

Totally agree on the rule ordering issue. That's often the silent killer.

We learned this the hard way - our fuzzy IP rule was merging profiles before our CRM ID sync could run, creating Frankenstein customers. The fix was to treat IP and device matches as "possible links" in a separate staging graph, not direct merges.

Your point about exact vs. fuzzy on categorical data is spot on. We even stopped fuzzy matching IPs entirely. Now it's either an exact match (treated as low confidence) or no match.


Trial first, ask later.


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Your observability instinct is spot on - you're basically correlating metrics with wildly different cardinality and expecting a clean output.

> shared or dynamic IPs (like from a corp

That's exactly the trigger. Corporate NATs and mobile carriers pool IPs. Your spike likely coincided with more mobile traffic or a remote work week. A single office IP hitting your threshold will merge everyone in that building who isn't logged in.

Treat IP and device matches like cheap reserved instances: useful for savings, but dangerous if you overcommit. Don't merge on them directly. Use them to build a low-confidence link table for manual review, not your primary graph.

Seen teams burn six figures on ad spend targeting these merged "super customers." Check your recent traffic sources for a spike in mobile or a new corporate partner domain.


- elle


   
ReplyQuote
(@gregoryt)
Reputable Member
Joined: 2 months ago
Posts: 418
 

Yeah, from a systems angle, that 0.8 threshold on IPs is like setting a super low CPU utilization alarm. It's gonna fire constantly and drown out the real signal. Your hunch about shared IPs is right.

I'm learning about CDPs myself, but in my Docker/Logging world, we treat IPs as high-cardinality labels, not merge keys. Couldn't the spike just be from one big corporate gateway getting reused?

A quick question: would adding a time window on that fuzzy rule help? Like, only merge IPs if they were seen within, say, 10 minutes of each other? Or is that not how the graph works?



   
ReplyQuote
(@briang)
Estimable Member
Joined: 3 months ago
Posts: 119
 

Good point about the time window. From what I've read, some CDPs do support that to reduce false merges. But wouldn't it still collapse all users from a shared office IP if they were active at the same time? Like during a lunch break when everyone's browsing.

I'm curious, in your logging setup, how do you handle that high cardinality without merging? Do you just keep IPs as separate tags and correlate them later?



   
ReplyQuote