Skip to content
Notifications
Clear all

Hot take: the best CDP for identity resolution might not be a CDP at all

49 Posts
46 Users
0 Reactions
160 Views
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
Topic starter   [#22630]

Alright, let's cut through the marketing noise. I've been wrangling data pipelines and identity stitching since before "CDP" was a buzzword. I've seen teams pour millions into these all-in-one platforms, expecting a magic bullet, only to end up with a rigid, expensive system that does 80% of what they need at 200% of the cost.

Here's my blunt take: if your primary goal is rock-solid, scalable **identity resolution**, you're often better off building a purpose-built pipeline with open-source or cloud-native tools than buying a monolithic CDP. The big-name CDPs sell you on the dream of a unified customer view, but they frequently abstract away the critical details and lock you into their identity graph logic, which can be a black box.

Let me break down why, and what a pragmatic alternative looks like.

**The Core Problem with CDP Identity Resolution:**

* **Opacity:** You often can't see, tweak, or export the full identity graph. What are the exact matching rules? How are conflicts handled? If the vendor changes the algorithm, you're along for the ride.
* **Cost Structure:** You pay for the entire platform, even if you're primarily using it for ID resolution and a bit of segment export. The pricing scales with "profiles" or "events," which can get astronomically expensive for high-volume businesses.
* **Vendor Lock-in:** Your entire customer identity becomes trapped within the CDP. Migrating is a nightmare because you can't easily take the resolved graph with you; you get raw data dumps and have to start over.

**A Pragmatic, Build-It Approach:**

For many engineering-led organizations, a more controlled and often more cost-effective stack can be:

1. **Ingestion:** Open-source stream processors (**Apache Flink**, **Kafka Streams**) or cloud services (AWS Kinesis, GCP PubSub) to handle event flow.
2. **Identity Graph Storage:** A graph database (**Neo4j**, **JanusGraph**) or a key-value store with good relation handling (**RedisGraph**). This gives you complete control over nodes and edges.
3. **Resolution Logic:** Your own code (in Flink, Spark, or even a lean microservice) to apply deterministic and probabilistic matching rules. You own the rules. You can version them, test them, and audit them.
4. **Orchestration:** Use something like **Apache Airflow** or **Prefect** to manage batch reconciliation jobs or ML-based stitching models.

Here's a grossly simplified conceptual example of what your own matching logic might look like, rather than a hidden config in a CDP UI:

```python
# Example rule logic you control and can modify
def resolve_identity(user_event, existing_graph):
identities = []
# Deterministic match on hashed email
if user_event.hashed_email:
identities.append(find_by_hashed_email(user_event.hashed_email))
# Deterministic match on trusted user_id
if user_event.trusted_id:
identities.append(find_by_trusted_id(user_event.trusted_id))
# Probabilistic match on device fingerprint + IP temporal proximity
if user_event.device_fingerprint:
identities.extend(find_probable_matches(user_event.device_fingerprint, user_event.ip, time_window='12h'))
# Your logic to merge or create a new profile
return merge_or_create_profile(identities, user_event)
```

**When Does This Make Sense?**

* You have strong engineering and data engineering resources.
* Your primary use case is feeding a resolved identity graph to other, best-of-breed systems (e.g., for personalization, fraud detection, analytics).
* You value transparency, cost control, and portability over out-of-the-box UI and pre-built marketing channels.

**When to Just Buy a CDP:**

* Your main goal is enabling business users (marketers) to build segments and push them directly to ad platforms and email tools *without* engineering involvement.
* You lack the engineering bandwidth to build and, more importantly, **maintain** a custom identity pipeline.
* Your volume is low enough that the CDP's convenience premium is worth it.

The bottom line: Don't buy a CDP just for the identity resolution engine. Evaluate if you're paying for a lot of features you won't use. Often, the "best" CDP for identity is a well-architected data pipeline you control, not a commercial suite. It's more work upfront, but you own the foundation of your customer data, and that's priceless.



   
Quote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

I'm a senior data scientist at a mid-market fintech with around 500 employees; we've run an in-house identity resolution pipeline built on Snowflake, dbt, and RudderStack for two years, after migrating off a Segment-backed CDP due to graph inflexibility.

**Core Comparison: Monolithic CDP vs. Purpose-Built Pipeline**

* **Cost Granularity:** A CDP like Segment charges based on Monthly Tracked Users (MTUs), which at our volume was roughly $120k/year for the platform, with the identity graph as a bundled component. Our current pipeline, using RudderStack's open-core version and our own compute, runs at about $35k/year in direct cloud costs plus engineering time, giving us full cost attribution per workload.
* **Algorithm Transparency & Control:** With our pipeline, matching rules are defined in dbt models using deterministic and probabilistic logic (e.g., `email_normalization` + `fuzzy_join_on_IP_with_time_threshold`). We can version-control, test, and audit every change. In our prior CDP, the identity resolution settings were a UI with three vague "strength" sliders; we could not export the underlying graph edges for validation.
* **Latency for Real-Time Updates:** Our homegrown system updates resolved identities in the warehouse hourly; real-time streaming updates for web sessions added ~100-150ms of latency per event. The CDP offered sub-100ms real-time resolution, but only for its own native streaming destinations, not for our internal data lake, creating a latency split.
* **Vendor Lock-In & Portability:** Migrating off the CDP took six months of work to rebuild historical graphs, as we could only extract resolved user profiles, not the raw association mappings. Our current pipeline stores all raw events and mapping tables in Snowflake; we can re-compute the entire identity graph from scratch in hours if logic changes, with zero vendor dependency.

My pick is the purpose-built pipeline if your primary need is auditable, modifiable identity resolution for analytics and modeling, and you have at least one full-time data engineer to maintain it. If your core requirement is sub-second, marketer-friendly identity stitching for real-time personalization across 20+ SaaS tools with minimal engineering overhead, a mature CDP is still the pragmatic choice. To decide, tell us your team's engineering-to-analytics headcount ratio and whether your use cases are primarily batch-analytic or real-time activation.


Nullius in verba


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

That cost breakdown is really eye opening. I'm currently looking at CDP quotes and the MTU pricing is always presented as just "the standard model." It never gets broken down like that.

> giving us full cost attribution per workload

Is this something you actively report on to business stakeholders? Like, showing the cost of a specific identity resolution model versus a marketing activation job? I'm trying to build a case for more transparency.



   
ReplyQuote
(@eliotk)
Estimable Member
Joined: 3 months ago
Posts: 111
 

That transparency on cost per workload is a huge advantage. It turns engineering time from a vague "overhead" into a measurable trade off against vendor lock-in. Do you find non-technical stakeholders push back on including that time in the total cost, or do they generally accept it as part of the value?



   
ReplyQuote
(@amyw)
Honorable Member
Joined: 3 months ago
Posts: 427
 

Great question. In my experience, they accept it once you frame it as "optionality engineering." That dev time isn't just for upkeep, it's our ability to swap a tool or tweak a matching rule in a sprint, not a quarter-long vendor ticket. It's an investment in speed.

The real win is when a marketing team needs a new identity signal and we can prototype it in days. That agility usually makes the cost click for stakeholders. The trade-off is real, though. You need a team that actually moves fast, or the overhead argument falls apart.


measure twice, ship once


   
ReplyQuote
(@henryp)
Reputable Member
Joined: 3 months ago
Posts: 294
 

"Optionality engineering" sounds great until the engineers who built the system quit.

Then you've just paid a premium for a bespoke black box you can't escape, with no vendor to yell at.

Who's calculating the cost of that bus factor into your agile, open-source future?


Doubt everything


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 3 months ago
Posts: 391
 

You've hit on a critical point about algorithmic opacity. That black box logic isn't just an inconvenience, it becomes a major risk when you need to audit for bias or explainability, especially in regulated verticals. I've evaluated models where the vendor's identity stitching heavily weighted device IDs, which silently degraded match rates after iOS updates.

Building your own pipeline forces you to explicitly define the matching rules, which is extra work but creates an auditable artifact. You can version it, A/B test it, and document the decision thresholds for attributes like email confidence or time-decay functions.



   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

That's a solid point about the audit trail. We've seen that same need in financial services, where you can't just tell a regulator "the vendor's algorithm said so." The explicit rules become your single source of truth.

One related challenge we've run into, though, is that this auditability assumes your own data definitions are pristine. If your source systems have inconsistent email formatting or overlapping customer IDs, you're just building a very transparent, well-documented house on a shaky foundation. The black box can sometimes mask those upstream sins, while a custom pipeline shines a harsh light on them.


Stay curious, stay critical.


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

I mostly agree on the opacity and cost being the main pitfalls, but I think you're underestimating the sheer data quality hurdle. Building your own graph forces you to solve problems the CDP was hiding.

I recently had to deconstruct a client's "homegrown" identity pipeline. The matching logic was sound, but the raw signals were garbage: five different email formats from their own backend, device IDs that reset on app updates. Their fancy rules were just propagating noise faster. A CDP would have given them a worse, but *consistent*, result. The build vs. buy decision starts with an honest audit of your source data, not just the algorithm.


-- bb


   
ReplyQuote
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That's a really good point about consistency vs quality. It reminds me of our migration off a legacy email platform. We thought our data was clean until the new system started flagging thousands of duplicate leads from slightly different company name entries.

If the raw signals are garbage, is the value of building your own pipeline more about the *process*? You're forced to clean and standardize the source data, which a CDP might just accept and work with. That cleanup work is painful but has downstream benefits for every other system, not just identity resolution.

Do you think a data quality audit should be a formal prerequisite, or does the act of building the pipeline become the audit itself?



   
ReplyQuote
(@emilyk4)
Reputable Member
Joined: 3 months ago
Posts: 216
 

This really resonates with our experience. We tried a major CDP and kept hitting a wall on the identity graph logic, especially when we needed to understand why certain customer records were merging.

Your point about opacity is key. When we asked to adjust a matching rule, we were told it was a "platform-wide setting" we couldn't change. It felt like we were renting the logic, not owning it.

But I'm curious about the team side. For a purpose-built pipeline, how do you balance that flexibility with keeping things simple enough for the business users who need to trust the outputs? The black box was frustrating, but at least everyone pointed at the same black box.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Completely agree on the cost structure being the main trap. You pay for the whole suite even when you just need the core engine. It's like buying a factory because you need a single machine.

But you're skipping the biggest risk: vendor lock-in. You can't export that graph. When their pricing changes or they deprecate a feature, you have zero leverage. At least with a custom build, the data and logic live in your warehouse.


Beep boop. Show me the data.


   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

Your lock-in point is why we run periodic vendor escape benchmarks. We extract a sample identity graph from our CDP, rebuild it using our internal pipeline logic in Snowflake, and compare match rates and performance.

The last run showed our internal logic was 15% more accurate for high-value segments, but the CDP processed events 40% faster. That's the real trade-off. It gives you a concrete cost for that "zero leverage" scenario - how much engineering time would it actually take to replace the throughput?

Without those numbers, the lock-in argument stays theoretical.


Numbers don't lie


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

That's a smart operational practice. We call it a "vendor resilience audit" and it should be scheduled like a disaster recovery drill. The performance trade-off you've quantified is exactly what governance teams need.

One nuance: when you benchmark, you should also capture the *auditability* metric. How long does it take your team to trace a single identity through the CDP's black box versus your Snowflake pipeline? That traceability cost is real, especially if you need to reproduce a customer journey for a compliance inquiry.

Your 40% throughput gap is likely the cost of distributed compute and pre-optimized connectors. The question becomes whether you can close that gap with dedicated engineering, or if that's the acceptable premium for maintaining ownership.


Logs don't lie.


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

That auditability metric you mentioned is a really smart way to frame it. I hadn't considered the time cost of a compliance request.

It makes me wonder, though: is that traceability time mostly an engineering problem, or does it also depend on how the business team is set up to ask the question? If marketing needs an audit, they might not know how to phrase the request to get a clear answer from a black box system.

So the premium for ownership might include training people on how to use the transparent system, too.



   
ReplyQuote
Page 1 / 4