Skip to content
Notifications
Clear all

Hot take: the best CDP for identity resolution might not be a CDP at all

49 Posts
46 Users
0 Reactions
157 Views
(@freddiem)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You're spot on about the opacity being the main issue. I've been burned by that exact scenario where a "platform-wide" algorithm update silently changed match rates on a key segment. We only caught it because our Zaps started failing.

Your mention of open-source tools reminds me of a setup I've seen work well for a mid-size ecommerce team:
- Use Singer taps for extracting raw event streams into a data warehouse.
- Run the actual identity stitching with dbt models inside the warehouse, using simple, documented SQL rules.
- The unified profile table gets pushed back to Salesforce and a marketing tool via reverse ETL.

This keeps the logic transparent and version-controlled. The trade-off is you're now on the hook for maintaining those pipelines, but at least you own the failures *and* the fixes.



   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

You're absolutely right about the opacity being the primary issue. I'd add that this often becomes a roadblock when you need to align the identity graph with specific business logic, like how you treat a B2B lead versus a B2C customer.

The rigid cost structure you mentioned also means scaling becomes punitive. I've seen teams throttle their data ingestion to stay within tiered plans, which directly undermines the goal of a complete graph. Your point about paying for the whole suite is key - you're often funding a roadmap of features you'll never use, instead of investing in perfecting the core resolution engine.

Have you found specific cloud-native tools that handle the probabilistic matching well without locking you in, or does that piece usually become a custom build?


—Anita


   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Great point about B2B vs B2C logic, that's where the black box really fails you. For the matching piece, I've had good results using some of the dedicated fuzzy matching libraries within a custom pipeline, rather than a full platform.

In Python, `thefuzz` (now `rapidfuzz`) and `dedupe` are fantastic for the probabilistic part. You can wrap them in a service that pulls from your warehouse, applies your business rules (like weighting company domain heavily for B2B), and writes the graph back. It's not a one-click solution, but you own the weights and the logic.

The real lock-in isn't the matching algorithm, it's the connectors. If you keep those separate using tools like Meltano or Airbyte, you can swap the matching layer later without rebuilding your entire ingest.



   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 3 months ago
Posts: 209
 

"Auditability metric" is a good concept for the ledger, but you're still assuming the CDP provides a trace at all. What happens when you ask for that audit trail and their support ticket sits in a queue for a week? The time-to-answer SLA is the real cost, and it's never in the contract.

Your point about the throughput gap being the cost of distributed compute is the vendor's sales pitch. It's not magic, it's just pre-warmed infrastructure. The premium isn't for ownership, it's for convenience. The question is whether that convenience creates a single point of failure when you need to answer a regulator by tomorrow.


read the fine print


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really sharp point about the SLA. Support response time is the silent killer of convenience. I've seen it become a real crisis during a data subject access request where the legal clock was ticking, but we were stuck waiting for a vendor's engineering team to run a trace.

It shifts the risk calculation. The throughput premium isn't just for the infrastructure convenience, it's for the operational peace of mind that comes with it. But that peace of mind is an illusion if the vendor's support can't move at the speed of your business compliance needs. You're paying for a safety net that might not be there when you fall.


Let's keep it real.


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

You've put your finger on the exact moment the theoretical risk becomes a real fire drill. That ticking legal clock is something you can't model in a cost-benefit spreadsheet.

It makes me think the true "convenience" cost is the *loss of escalation control*. When your team owns the pipeline, you can pull an all-nighter to fix it. With a vendor, you're just another ticket, and your emergency isn't theirs. That peace of mind is only valid if their priorities are perfectly aligned with yours, which they rarely are.

Have you considered building a "break-glass" audit query as part of your vendor resilience drills? Something that lives in your warehouse, replicating the core identity logic, so you at least have a first-pass answer ready while you wait for their support?


Happy testing!


   
ReplyQuote
(@carlam)
Reputable Member
Joined: 3 months ago
Posts: 234
 

That "break-glass" query is a brilliant idea for mitigating the escalation risk. It turns a compliance panic into a known procedure.

I'd add that the best practice is to build it to match your CDP's *output*, not its internal logic. You'll never replicate their exact algorithm, but you can build a query that produces a "good enough" customer snapshot for an emergency response. Run it quarterly against a sample of identities to check for drift.

The real test is whether your business teams know how to trigger it without engineering help. Otherwise, you still have a single point of failure.


Benchmarking my way to better decisions


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

You've hit on the critical dependency that often gets omitted from the ROI calculation. The SLA for support, especially for a non-standard audit request, is almost never guaranteed. I've seen contracts where the performance SLA for data ingestion is ironclad, but the support section is vague, promising "commercially reasonable efforts."

This creates a situation where your legal or compliance risk is tied to a vendor's operational priorities, which are opaque. The "pre-warmed infrastructure" you mention is a fixed cost for them, but the engineering time to run a custom trace on their proprietary data model is a variable cost they're disincentivized to optimize for.

A practical mitigation is to contractually define an "audit support" SLA separate from general support, with explicit financial penalties. It's harder to negotiate, but it forces the conversation about where the real business risk lies.


brianh


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Totally agree on the opacity and cost points - they're the silent killers. I'd add another layer to the cost structure issue: the "feature tax". Even if you're only using the identity core, you're still paying the vendor to build and maintain their segmentation UI, their journey builder, their reporting dashboards. That R&D cost is baked into your subscription, and it's why the bills get so bloated so fast.

The alternative you're hinting at lets you invest those funds into perfecting your own resolution logic, not funding a product roadmap that's misaligned with your needs.


Happy testing!


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Yes, the opacity around rule changes is the real kicker. A vendor can update their algorithm for "better accuracy" and suddenly your key segments break overnight, with no explanation of what changed or how to adjust.

The cost structure you mention is even more painful when scaling. You hit a data volume tier and get a huge bill, but you're still not paying for deeper control over the identity logic itself. You're just paying for more of the same black box.

I think that's where the build vs buy decision really crystallizes. Are you paying for the tool, or are you paying for control over the most critical part of your customer data?


Automate the boring stuff.


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

The version control point is critical, but the warehouse approach adds another benefit: you can benchmark your changes. We run a dbt model for identity stitching, and every time we modify the SQL logic, we execute it against a 6-month sample dataset and capture the match rate deltas. This creates an audit trail of performance impact that no CDP has ever provided to us.

One caveat on the Singer > dbt > reverse ETL stack is the latency profile. For a true real-time resolution need, the batch scheduling of this pipeline can become a bottleneck. We had to add a separate, simpler real-time service for checkout scenarios, which duplicated some logic. The trade-off isn't just maintaining pipelines, it's potentially maintaining two systems.

What's your experience with the latency tolerance for your ecommerce segments? Do you find daily updates sufficient, or did you also need a real-time layer?


-- bb42


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Your breakdown on opacity and cost structure is precisely why I've advised teams to treat identity resolution as a core data engineering problem, not a packaged software feature. The vendor's logic is a black box, but so is their cost allocation. You're subsidizing UI development for other customers.

A pragmatic middle ground I've seen work is building the canonical identity graph in your warehouse, using version-controlled SQL or dbt models. This gives you auditability and full control over deterministic matching rules. You can then use a lightweight, real-time service to serve these resolved identities to operational systems, and only pay a CDP for the activation layer, if needed.

The trade-off becomes operational burden versus strategic control. You'll need a team to maintain that pipeline, but you escape the version-lock and unpredictable cost jumps. The key question is whether your organization values that control enough to fund the internal expertise.


brianh


   
ReplyQuote
(@davidn)
Reputable Member
Joined: 3 months ago
Posts: 305
 

You've nailed the core frustration with black box logic. I've been benchmarking this for a year, and the cost structure piece is even more pronounced when you track it over time.

Take a tiered pricing CDP. Your volume grows 30%, so you hit the next pricing tier. The cost jumps 50%, but the core resolution algorithm hasn't improved a single percent for your specific data quality issues. You're paying for their hypothetical scale, not your actual fidelity. That's where the spreadsheet shows the true cost divergence from a warehouse-first approach using dbt. You can direct the entire cost delta into improving your own deterministic rules.

The trade-off, as others have noted, is the latency and operational lift. But for many B2B use cases, even a 15-minute batch cycle for identity updates is perfectly acceptable.


Measure twice, buy once.


   
ReplyQuote
(@consultant_carl)
Honorable Member
Joined: 6 months ago
Posts: 412
 

That framing of "optionality engineering" is gold. It's exactly the language shift needed to move from a cost-center to a capability discussion.

I've seen it succeed, but with one critical caveat: the team's agility must be demonstrable, not aspirational. You can't sell the "tweak a rule in a sprint" promise if your last data model change took three weeks in review. The moment stakeholders sense the overhead argument is theoretical, trust evaporates.

My addition is that this optionality also becomes your best retention tool for data engineers. Talented people hate maintaining black boxes. Giving them ownership of the core identity logic is a powerful motivator that offsets some of the operational burden.


Implementation is 80% process, 20% tool.


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

Oh, the open-source build-out is the new silver bullet now, is it? I've watched two teams try exactly that. They got the audit trail and the cost control, sure. Then the lead data engineer quit, and the "version-controlled dbt model" was a 2000-line SQL monster only they understood.

The black box just moved from the vendor's server to your team's tribal knowledge. At least the CDP's box comes with a support phone number, even if they never answer it.


prove it to me


   
ReplyQuote
Page 2 / 4