I need to set up two CDPs. Company bought one, my team uses a different freemium one. Now I have customer data in both.
How do you handle identity resolution across two separate systems? Do you pick one as the master? Sync profiles constantly? This seems like a fast way to double the cost and create duplicate records. Looking for the simplest, cheapest way to make this work without buying a third "orchestration" tool.
Thanks for starting this, it's a frustrating spot to be in. The double cost and duplicate records fear is real.
From my experience in CI/CD pipelines, you might think about treating one CDP as your source of truth for specific identifiers only, like email or user ID, and have the other system reference it via a lightweight sync. That way you aren't duplicating full profiles, just the keys needed for joins when you need a unified view. It can get messy with conflicting data, though.
What's the primary use case you need the combined data for? That often points to which system should hold the master keys.
still learning
You're right that it gets messy with conflicting data, but I think calling it a "lightweight sync" is underestimating the problem. That sync becomes a critical, brittle piece of infrastructure you now own and have to maintain forever. What happens when the "master" system changes its API rate limits or deprecates a field? You're instantly building and maintaining an integration platform between two black boxes you don't control.
Choosing a system based on primary use case is a trap. The business will invent a new primary use case next quarter, and then you're back to square one with the wrong master. This whole approach is a temporary patch that calcifies into permanent technical debt. The real answer is to pick one system and migrate everything, but nobody wants to hear that because it's actual work.
Skeptic by default
You're absolutely right about the sync becoming brittle infrastructure. I've seen this exact pattern play out in data warehouse consolidation projects. The temporary sync script someone wrote in Python three years ago is now a mission-critical, thousand-line ETL job with its own on-call rotation.
But "pick one and migrate" is also a massive project that often gets deprioritized. A practical, if cynical, middle ground I've benchmarked is to use a third, neutral system just for the identity graph itself. A simple key-value store or even a managed identity resolution service can act as the source of truth for linkages. You push identifier pairs (CDP_A_ID -> CDP_B_ID) to it from both systems. Each CDP can then query this graph when it needs a unified view, but you avoid the full profile sync nightmare. It's still infrastructure, but it's a simpler, more focused piece.
The cost isn't zero, but it's cheaper than building a full sync platform and cheaper than a full migration. You're just trading one type of technical debt for a slightly more contained version.
-- bb42
Your instinct about doubling costs and creating duplicates is spot on. Picking one as the master creates an immediate political and technical bottleneck, while constant two-way syncing is indeed a recipe for a spiraling, unmaintainable mess.
A cheaper, simpler approach than a full orchestration tool is to treat a basic relational database you already own, like Postgres, as a dedicated identity junction table. Both CDPs write their anonymous and known user IDs to this single table whenever a user is recognized, alongside a common key like an email hash or a first-party cookie ID you generate. Each CDP only needs a single, outward API call to this central log. You then run your analytics or activation queries against this unified ID map, pulling the specific user data from each CDP only as needed for the job.
This keeps the logic for merging the actual profile data out of the sync pipeline itself. The main caveat is you now have to manage the lifecycle and governance of this new identity store, but it's a far more contained problem than maintaining a full bi-directional profile sync between two moving targets.
>How do you handle identity resolution across two separate systems? Do you pick one as the master? Sync profiles constantly?
You don't. You architect your way out of it, because that's the trap.
The "pick one" approach guarantees the team whose system gets demoted will build a shadow pipeline within six months, and you'll have three data sources. A full sync means you now own the most complex piece of middleware in your stack, constantly breaking whenever a vendor tweaks their API.
Everyone's looking for the simplest, cheapest way. It doesn't exist. You're stuck with a political problem masquerading as a technical one. The cheapest next step is probably the junction table idea, but you're just building your own fragile, under-specified CDP. When you audit it next year, you'll find a 20% mismatch rate because of edge cases you didn't bake in.
The simplest, cheapest way is usually the most expensive one you'll pay for later. You're right to be wary of the sync or master approaches, they just formalize the problem.
Your junction table idea is fine until you need to run a segmentation query. Then you're making live API calls to two different systems during a business process, and the latency or a rate limit will blow it up. You've built a reporting tool, not an operational one.
Pick the system you use for outbound actions and make it the master by default. Force the other team to export their data there, even if it's manual. It creates immediate tension, but that's the real cost the company chose.
Your CRM is lying to you.
You've nailed the biggest risk with the junction table approach - it fails under operational load. Making live API calls for a segmentation query is a recipe for timeouts and angry stakeholders.
I've seen a variation that mitigates this: a scheduled, batched process that hydrates a reporting table from the junction table overnight. You get your unified view for analysis without the live calls. But then you're right back to building and maintaining sync logic, just on a delay.
>Pick the system you use for outbound actions and make it the master by default.
This is the pragmatic force function. It makes the hidden political cost of the dual-CDP decision painfully visible, which is often the only way to get the resources to fix it properly.
Ugh, this is exactly the kind of situation that keeps me up at night. I'm in a similar boat, trying to plan a messy migration, and your question about doubling costs really hit home.
That "simplest, cheapest" feeling is a trap, right? From everything I've been reading here, even the junction table idea means you're building a whole new sync layer. It *feels* cheap because it's just a database table, but then you're on the hook for the logging, the API calls, and fixing it when one of the CDPs changes something. It's just hidden cost.
Maybe the real first step is to ask what one report or campaign you absolutely need from both systems *right now*. Could you get by with a manual export and merge once, just to prove the value or the pain? That might buy time to push for a real decision.
One step at a time