Hey folks, hoping to get some collective wisdom here. I’m helping our marketing team troubleshoot our CDP's identity resolution, and we're seeing a huge number of merged customer profiles that just don't seem right. Our "identity graph" is lumping together what are clearly different people, which is throwing off our segment counts and activation.
We’re using a deterministic rule set based on email and user ID, with some fuzzy matching on IP addresses and device IDs for unknown visitors. The volume of merged profiles has spiked about 300% in the last month. 😬
Here’s a simplified version of the rule configuration we’re using:
```json
{
"identity_rules": [
{
"priority": 1,
"rule_type": "deterministic",
"attributes": ["email", "user_id"],
"action": "merge"
},
{
"priority": 2,
"rule_type": "fuzzy",
"attributes": ["ip_address", "device_id"],
"threshold": 0.8,
"action": "merge"
}
]
}
```
From my observability world, this feels like an overly sensitive correlation rule creating noise. I’ve been checking the raw events feeding the graph, and I'm wondering:
* Could shared or dynamic IPs (like from a corporate network or mobile carrier) be causing the fuzzy rule to over-match?
* Is there a common pitfall in how device IDs are being generated or persisted across different platforms (web vs. app)?
* Should we be adding more negative correlation rules to prevent merges when certain high-confidence fields (like last name) don't match?
I’d love to hear from anyone who has tuned these rules before. What metrics did you track to validate your merges? Did you have to dial back fuzzy matching significantly? Any dashboard screenshots of your merge rates over time would be super helpful!
Dashboards or it didn't happen.
Your fuzzy rule threshold at 0.8 is way too aggressive for IP and device ID. That's almost certainly the culprit.
Shared corporate or public IPs, plus mobile devices that share ad IDs, will hit that similarity score easily. It'll merge whole office buildings or families into single profiles.
Drop the fuzzy threshold to at least 0.95 and check your match keys. You likely have a data quality issue where placeholder or null device IDs are being scored as identical.
Prove it with a benchmark.
user1054's point about the fuzzy threshold is correct, but the 0.95 suggestion could still be problematic. The core issue is using fuzzy matching on inherently unstable identifiers for deterministic merging.
An IP address isn't a fuzzy identifier, it's a categorical one. Two users sharing a corporate NAT have identical IPs, not similar ones. Fuzzy logic on a full IP string gives you a false sense of precision. A better approach is to separate the logic: exact match on a hashed identifier, then apply a completely different rule for shared-network scenarios based on a confidence weight, not a similarity score. Device IDs should be handled the same way - either they match exactly or they're different.
This often points to a rule ordering problem. If your deterministic email/user_id rule runs first and creates a profile, then your fuzzy IP rule runs and merges another profile onto it, you've now contaminated your graph. You need to quarantine matches from low-fidelity signals, not merge them directly.
brianh
Totally agree on the rule ordering issue. That's often the silent killer.
We learned this the hard way - our fuzzy IP rule was merging profiles before our CRM ID sync could run, creating Frankenstein customers. The fix was to treat IP and device matches as "possible links" in a separate staging graph, not direct merges.
Your point about exact vs. fuzzy on categorical data is spot on. We even stopped fuzzy matching IPs entirely. Now it's either an exact match (treated as low confidence) or no match.
Trial first, ask later.
Your observability instinct is spot on - you're basically correlating metrics with wildly different cardinality and expecting a clean output.
> shared or dynamic IPs (like from a corp
That's exactly the trigger. Corporate NATs and mobile carriers pool IPs. Your spike likely coincided with more mobile traffic or a remote work week. A single office IP hitting your threshold will merge everyone in that building who isn't logged in.
Treat IP and device matches like cheap reserved instances: useful for savings, but dangerous if you overcommit. Don't merge on them directly. Use them to build a low-confidence link table for manual review, not your primary graph.
Seen teams burn six figures on ad spend targeting these merged "super customers." Check your recent traffic sources for a spike in mobile or a new corporate partner domain.
- elle
Yeah, from a systems angle, that 0.8 threshold on IPs is like setting a super low CPU utilization alarm. It's gonna fire constantly and drown out the real signal. Your hunch about shared IPs is right.
I'm learning about CDPs myself, but in my Docker/Logging world, we treat IPs as high-cardinality labels, not merge keys. Couldn't the spike just be from one big corporate gateway getting reused?
A quick question: would adding a time window on that fuzzy rule help? Like, only merge IPs if they were seen within, say, 10 minutes of each other? Or is that not how the graph works?
Good point about the time window. From what I've read, some CDPs do support that to reduce false merges. But wouldn't it still collapse all users from a shared office IP if they were active at the same time? Like during a lunch break when everyone's browsing.
I'm curious, in your logging setup, how do you handle that high cardinality without merging? Do you just keep IPs as separate tags and correlate them later?