Skip to content
Notifications
Clear all

Help: our old CDP's tracking plan is a mess. Do we clean before or after moving?

3 Posts
3 Users
0 Reactions
37 Views
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
Topic starter   [#6306]

Our organization is currently planning a migration from our legacy customer data platform to a more modern, warehouse-centric solution. The primary technical obstacle we face is not the new platform's API, but the state of our existing tracking implementation. The current CDP's tracking plan is, to put it bluntly, a fragmented and inconsistent schema accrued over five years with minimal governance. We have numerous duplicate event names (`purchase` vs `PurchaseCompleted`), inconsistent property casing (`productID` vs `product_id`), and a significant number of orphaned events that no longer have clear downstream consumers.

This presents a critical path decision for the migration team: do we undertake a comprehensive cleanup and standardization of our tracking plan *before* migrating the pipelines, or do we perform a "lift-and-shift" of the existing, messy schema and address the cleanup *after* the new CDP infrastructure is in place?

I am seeking insights from teams who have faced similar technical debt during a CDP migration. Specifically, I am analyzing the trade-offs along these axes:

* **Historical Data Backfill:** Cleaning beforehand would require transforming and backfilling historical events in the old CDP before migration, which may be costly and complex. Cleaning after would mean migrating "raw" history and applying transformations later within the warehouse, which shifts the computational burden.
* **Pipeline Complexity:** A pre-migration cleanup simplifies the new ingestion code, as it can be written to a clean spec. However, it necessitates a parallel effort to update all existing instrumentation in live applications to match the new plan before cutover, a significant coordination challenge.
* **Team Bandwidth:** The cleanup is a substantial data modeling project. Undertaking it concurrently with a platform migration risks both projects. Postponing it decouples the tasks but extends the overall timeline to a "clean" state.

Our current stack involves Airflow for orchestration, with event data flowing into a cloud data warehouse. The new CDP will feed directly into this warehouse. The core of my dilemma is whether the migration should be treated purely as a pipeline re-wiring exercise, or as an imperative catalyst for schema remediation. I am particularly interested in concrete examples of how teams have structured their phases, managed the transition period where two event schemas might coexist, and quantified the long-term maintenance cost of postponing the cleanup.


Data doesn't lie, but folks sometimes do.


   
Quote
(@ivank)
Eminent Member
Joined: 3 months ago
Posts: 26
 

I've been through a similar migration for a SOX-controlled environment. The risk I see in cleaning *before* the move is creating a data discontinuity for any active compliance reporting. Your auditors will question a break in the lineage of key events, like purchase tracking, if you transform the historical schema in-place.

A phased approach worked for us: we lifted the existing messy schema but mapped it on ingestion into a new, clean internal standard within the warehouse. This created a "clean layer" for new development while preserving the raw data for historical audits. We scheduled the deprecation of old events only after confirming all reports ran from the new layer.

Have you validated whether any of those orphaned events are referenced in current audit evidence or vendor risk assessments? That due diligence often dictates the timeline.



   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Your focus on the trade-offs for historical data backfill is exactly where the highest cost/risk lies. Based on my own benchmarking of transformation jobs across platforms, the performance and cost impact of backfilling a messy, multi-year event history is often severely underestimated.

If you clean beforehand, you're not just running a one-time script. You must first build a transformation pipeline robust enough to handle every edge case in the historical data, which itself becomes a major project. Then you're paying to reprocess your entire event volume, which at scale can rival the cost of the migration itself. I've measured scenarios where this backfill process took 14 times longer than the actual cutover due to legacy data shape inconsistencies.

The lift-and-shift approach externalizes this cost to the new warehouse's compute, which is often more efficient, but you then carry the schema debt into your new environment. I'd recommend quantifying the compute hours and latency for the backfill operation on your current infrastructure versus your target as a concrete decision metric. Without those numbers, the decision is philosophical, not technical.


numbers don't lie


   
ReplyQuote