Skip to content
Notifications
Clear all

Guide: how to audit a CDP's identity resolution quality with your own data

1 Posts
1 Users
0 Reactions
0 Views
(@kittycat)
Trusted Member
Joined: 1 week ago
Posts: 31
Topic starter   [#8691]

So, you’re thinking about buying a Customer Data Platform and the vendor is promising 90%+ identity resolution rates. Sounds great, right? But here’s the thing: that number is almost meaningless without context. It could be based on pristine, first-party data from a perfect test case. The only way to know how it performs for *you* is to run your own data through it.

I’m a big believer in treating this like an A/B test. You need a controlled, statistically sound method to audit the match rate you’ll actually get. The good news is, you can set this up yourself without a full implementation.

Here’s the pragmatic approach we used last year when evaluating three platforms:

**First, isolate your test population.** Take a sample of your known users (say, 50k logged-in customers) and hash their identifiers (email, phone, etc.). This is your "truth set." Then, strip away those known IDs and see what the CDP can stitch back together using just their anonymous tracks (like device IDs, cookie IDs, IP). The delta between what you know and what they find is your true incremental match rate.

**Second, measure consistency, not just volume.** A high match rate is useless if it's unstable. Ask for the "overlap rate" between daily resolution graphs. If User X is recognized as one profile on Monday but a different one on Tuesday, that's a huge red flag for activation campaigns. We found one platform with a 75% match rate had a daily consistency of under 40%—basically useless for personalization.

**Finally, audit the "why."** When they give you a match, demand the logic path. Was it a deterministic email match? A probabilistic device graph? A fuzzy name match? You need to know this to gauge scalability and privacy compliance. We built a simple tracking sheet to categorize match types, which revealed that one vendor's "high-quality" rate was heavily reliant on a third-party graph we weren't comfortable using.

The goal isn't to achieve a perfect score, but to understand the composition and stability of the matches. A platform with a lower, but 95% consistent, deterministic rate is often better for conversion work than a high, volatile probabilistic one.

Has anyone else run a similar audit? I’d love to compare methodologies on how you tracked consistency over time.

—kc


Sample size matters.


   
Quote