The company stage point is real, but I've seen that "messy, multi-brand data" scenario play out in two ways.
One team went all-in on a CDP, then spent 18 months building internal tooling anyway just to validate its merges and answer basic business questions about user journeys. They paid for two solutions.
The other team started with a simple deterministic match in their data warehouse, instrumented every merge decision with lineage tags, and only scaled complexity when the business could articulate the cost of a false match. Their "overhead" became a documented data asset.
A vendor SLA doesn't fix a missing spec. The question isn't just what phase you're in, but whether you're willing to treat identity logic as critical infrastructure or outsource the blueprint.
shift left or go home
Your second example captures the core engineering principle at play: observability and control as a first-class output, not a happy accident. Instrumenting every merge decision with lineage tags is essentially building a real-time audit log, which is something a CDP often treats as a premium feature or an afterthought.
I've seen teams take that a step further by treating those lineage tags as a source for a service-level objective on match quality. They expose a metric like `identity_graph_match_confidence`, derived from their own rule weights, and set an SLO against it. If a new data source or rule degrades the score, it fails a canary check in the pipeline. This turns abstract "health" into something you can alert on and defend in a postmortem.
The cost of that instrumentation is real, but it's a fixed cost you amortize over every future audit, investigation, and model retraining. Paying a CDP while also building internal tooling is the worst of both worlds.
Data over dogma
Yeah, the lock-in to a black-box identity graph is exactly what pushes teams toward GitOps for their data pipelines. If your core stitching logic is code in a repo, you can diff, rollback, and see exactly what changed. No more surprise algorithm updates from a vendor.
I've seen teams implement this with Argo CD or even just GitHub Actions. Your deterministic matching rules live in config files alongside your application code. A merge to main triggers a pipeline that runs the new logic in a canary mode first, comparing results to the old graph. It turns a scary CDP update into a reviewed PR.
That way, your identity resolution has the same audit trail and control as your deploys. You own the spec.
git push and pray
That point about paying for control hits directly on the licensing model. You're often paying per profile or event, which is a tax on data volume, not a fee for the intellectual property of the matching logic itself. The vendor's incentive is to keep that logic proprietary.
A related caveat: even if you accept the black box, the lack of a true change log makes regression analysis nearly impossible. When a segment breaks, you can't bisect through historical rule versions to pinpoint the change. You're left comparing static snapshots, which is a poor substitute for understanding causality. This turns every "improvement" into a potential production incident without a clear root cause.
Data is the new oil – but only if refined