Skip to content
Notifications
Clear all

Unpopular opinion: sometimes it's cheaper to keep the old CDP running for archives.

6 Posts
6 Users
0 Reactions
22 Views
(@billyp)
Reputable Member
Joined: 3 months ago
Posts: 284
Topic starter   [#27888]

Okay, hear me out. We just finished a massive migration from Mixpanel to a newer CDP, and the biggest surprise win was a decision we made halfway through: **we kept the old instance alive, read-only, for historical queries.**

I know, I know. It sounds like paying for two houses. But when we ran the numbers, the cost to backfill *all* historical events (we're talking 3+ years, billions of events) into the new systemβ€”both in engineering hours and the new platform's data storage feesβ€”was astronomical. Like, "hire two more analysts for a year" astronomical.

We realized our historical data needs fell into two buckets:
1. **Trend analysis & YoY comparisons:** Rare, but crucial for board reports.
2. **Ad-hoc user journey deep-dives:** "What did this segment do back in 2021?"

For #1, we pre-aggregated the key metrics we knew we'd need into our data warehouse during the migration. For #2, we simply kept the old Mixpanel project on a frozen plan. It's now just another data source in our Looker, accessed only when needed.

The cost? Less than 15% of the backfill quote. The engineering lift? Minimal. No wrestling with ancient, undocumented schema quirks in the new system.

Sometimes the "clean break" migration isn't the most efficient. If your old CDP has a reasonable archive fee, consider:

* **Leaving it as a read-only reference system** for <2% of queries.
* **Aggregating key historical metrics** you *know* you'll need into your warehouse.
* **Redirecting only net-new events and active user profiles** to the new platform.

This hybrid approach saved our timeline and budget. Anyone else done something similar, or am I just justifying our madness? 😅

Billy


Always A/B test.


   
Quote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

That's not an unpopular opinion, it's a pragmatic one. We did the same moving from a legacy marketing automation platform. The backfill cost and schema translation effort was insane.

The key is making that old instance truly read-only and cutting off any new data flows. Then it's just a static archive, not a live system. We put it behind a single, documented API endpoint so analysts don't even need direct access.

Where this gets messy is when legal asks for a user data purge. You now have to coordinate a delete across two systems. Make sure your retention policies are aligned.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

You're so right about the data purge headache! That's the hidden cost a lot of teams forget to budget for.

We set up a scheduled, quarterly "reconciliation audit" specifically for this. It runs a script to check a sample of IDs purged from the new system against the old archive, just to make sure nothing slips through. It's a bit of extra process, but it keeps legal and compliance happy.

And that single API endpoint idea is brilliant - it saves so much onboarding time for new team members who just need to pull old trends.


null


   
ReplyQuote
(@chloep)
Reputable Member
Joined: 3 months ago
Posts: 292
 

Totally, that "paying for two houses" feeling is exactly why most orgs don't even run the numbers. They just assume the backfill is a mandatory tax for the migration.

Your breakdown into those two buckets is spot on. The real question everyone should ask is: what's the *query latency requirement* for the archive? If the answer for those deep dives is "a few minutes is fine," then a frozen instance is perfect.

Where I've seen this go sideways is when the old platform decides to triple the price of their "legacy storage tier" after 18 months. Suddenly that 15% cost balloons. Did you negotiate a fixed, long term rate with Mixpanel before pulling the trigger on keeping it frozen, or are you just hoping their pricing page stays friendly? 😬


Demos are just theater. Show me the real workflow.


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

That point about the old platform tripling prices later is a real fear. We kept our old Segment space running and just got our first "infrastructure review" email, which I'm pretty sure is the first step toward a price hike.

> what's the query latency requirement for the archive?

This question is so important. We found even our "urgent" historical questions could wait an hour. So we actually shut down the old instance completely and just keep snapshot .parquet files in cold storage. If we need something, we spin up a cheap query engine to look at it. Takes longer, but costs almost nothing. Have you seen that approach work, or is restarting a service like that too much of a hassle?


null


   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 4 months ago
Posts: 723
 

Latency requirement is the right question. We benchmarked this.

The cost of a 'warm' read-only instance vs spinning up a query engine for parquet files depends entirely on query frequency. If you're doing more than a few dozen complex historical queries a month, keeping it warm is cheaper than the engineer-hours wasted babysitting spin-ups.

We didn't negotiate a fixed rate. Our risk mitigation was a monthly export of the archive to blob storage, ready to load into a warehouse if Mixpanel got predatory. The egress cost is the kill switch.


Benchmarks don't lie.


   
ReplyQuote