Skip to content
Notifications
Clear all

Hot take: migrating your CDP is a better time to fix your data model than any

7 Posts
7 Users
0 Reactions
4 Views
(@aidenf)
Reputable Member
Joined: 3 months ago
Posts: 219
Topic starter   [#28620]

Okay, hear me out. We all know our data models aren't perfect. There's always that one property that should have been an array, or that custom event name that's inconsistent, or the legacy user table that never got merged. But the thought of fixing it *within* your current CDP, with all the live pipelines and active dashboards? Terrifying. It's like trying to rebuild the engine while the car's going 60 down the highway.

Migrating from one CDP to another is that rare, golden opportunity. You're already planning to move the data. Why move the *mess*? You get to:
- **Redefine your core entities** (Users, Accounts, Products) with clean schemas from day one.
- **Standardize event naming** (we finally settled on `verb_noun` like `pricing_page_viewed`).
- **Backfill only the clean, transformed historical data** you actually need for models.
- **Re-wire downstream tools** (like your sales enablement or conversation intelligence platforms) to the new, sane data model, not the old patchwork.

We just finished moving from CDP A to B, and the best decision was using the migration project to finally fix our lead scoring logic. Instead of trying to translate our old, convoluted `lead_score` property (which was calculated three different ways over the years), we rebuilt the logic in the new CDP using a unified set of events and traits. The sales team is now getting scores that actually make sense.

Has anyone else used their CDP migration as a "data model reset" button? What was the biggest schema flaw you finally got to fix?

— Aiden


Let the machines do the grunt work


   
Quote
(@heidir33)
Reputable Member
Joined: 2 months ago
Posts: 270
 

This is such a great point. The line about rebuilding the engine at 60 mph is exactly right, and it's something we've struggled with for months.

When you say you redefined your core entities with clean schemas, did you do that before you started the actual data migration, or was it an iterative process during the extract/transform? I'm asking because we're in the early planning stages for a similar move, and I'm worried about locking in a new model too early before we fully understand how the downstream tools will need to consume it.

Also, curious about the backfill process - how did you decide what historical data was "actually needed" versus nice to have? That seems like a massive point of potential scope creep.



   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

We did the schema work upfront, but with an escape hatch. We built a prototype transformation pipeline in Datadog's data streams that mirrored our planned new model and ran it on a week of live data. This gave us concrete, flawed outputs to show the downstream teams (finance, product) for validation before we locked anything in. It's cheaper to throw away a week's worth of test transforms than to commit to a full migration.

For backfill, we used a simple rule: only data referenced by a dashboard or alert in the last 90 days was "actually needed." Everything else was archived to cold storage with a known restore path. This cut the historical load by about 70% and forced us to confront which legacy metrics were truly business-critical.


null


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Absolutely. The mental model of "rebuilding the engine at 60 mph" is precisely why these migrations are such a strategic inflection point. You've correctly identified the core benefit: you aren't just shifting infrastructure, you're forcing a re-evaluation of the data contract.

Your point about **re-wiring downstream tools** is critical and often underestimated. In our migration from Segment to Rudderstack, we documented every single destination and the specific fields it consumed. This exposed that 40% of the fields being sent were artifacts from old A/B tests or features that no longer existed. The migration wasn't just about moving data; it was about formally deprecating those fields and updating the API contracts with tools like Braze and Salesforce. The technical debt wasn't in the CDP, it was in the implicit, undocumented dependencies.

However, a caveat from our post-mortem: this approach can create a "big bang" risk if your new model is too ambitious. We phased it, fixing the egregious schema issues and standardizing events in the migration, but we deliberately postponed a complete "core entity" overhaul. We treated that as a separate, subsequent project fed by the now-cleaner data. The migration gave us the foundation, not the final architecture.


No free lunch in cloud.


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Spot on about fixing logic during the migration. We did the same with our alerting thresholds, which had been patched over for years.

Instead of porting the old spaghetti logic, we rebuilt the alert rules using the clean event streams, which finally let us implement proper severity levels and fatigue delays. The migration gave us the political cover to say "the new system needs new rules" and deprecate the fragile ones.

Be ready for the lag, though. Our product team took about a month to fully trust the new lead scoring alerts, even with side-by-side comparisons. You'll need to run dual-wiring for a bit.


Sleep is for the weak


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 3 months ago
Posts: 668
 

Love that approach with the prototype pipeline. We did something really similar with a small Lambda that transformed and wrote sample events to a test bucket in our new structure.

Your 90-day rule for dashboard references is genius. We ended up doing something a bit softer - any alert that had triggered in the last quarter got its data marked for migration. It caught a few dormant-but-critical business health checks that hadn't been looked at recently, but were still referenced in runbooks.

How did your finance team react to the concrete, flawed outputs? Ours had a few "oh, that's not right" moments that saved us from a major modeling error early on.


cost first, then scale


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

That prototype validation phase is critical. Our finance team had the same reaction, mostly because they were looking at aggregate reports that hid the underlying junk. Seeing raw events flagged a dozen different "total_amount" fields from years of payment system iterations. They thought it was one field.

We still got burned, though. The prototype used a week of recent data, which was too clean. The real backfill hit ancient event batches with null values in fields we'd assumed were mandatory. The new model choked. So our "escape hatch" turned into a six-week detour writing conditional logic for data we didn't even want to keep. Sometimes you can't win.

Your alert-trigger rule is smart. It surfaces the zombie processes that everyone forgets about until they wake up.


null


   
ReplyQuote