A common pain point during CDP migration is the data integrity gap that emerges between the old and new systems. Relying solely on a "big switch" date often leads to discrepancies in user analytics and model training, as historical context in the old CDP is severed from new events in the new one. A more financially and operationally sound approach is to run both systems in parallel using a feature flag.
This can be implemented by abstracting your event emission logic. Instead of calling the CDP's SDK directly, route all events through a central function or service. This function consults a feature flag configuration—which can be user, session, or account-based—and decides on the routing. The key patterns are:
* **Dual-write:** Send a copy of every event to both CDPs. This ensures complete parity but doubles your egress/event volume costs during the transition. Monitor this period closely.
* **Progressive cutover:** Use the flag to send a percentage of traffic (e.g., 10%) to the new CDP, validating its accuracy before increasing the load. This controls risk and allows for cost comparison on a like-for-like basis.
The configuration must be dynamic and controllable without a code deploy. For example, using LaunchDarkly, Flagsmith, or even a managed environment variable service. The logic should be simple and fail-open to your primary CDP to avoid data loss.
Critical considerations for this phase:
* **Cost:** You will incur double costs for the duration of the dual-write. Calculate this upfront and treat it as a necessary migration budget item.
* **Schema Alignment:** Ensure the event structure (properties, naming) is compatible with both destinations. This often requires a translation layer or configuring the new CDP to accept the existing schema.
* **Downstream Impact:** Analytics dashboards and data pipelines connected to the old CDP will now receive only a portion of events if using progressive cutover. They must be reconfigured to source from the new CDP or a merged stream to remain accurate during the transition.
The final step is to remove the flag and the old CDP integration once you've validated data consistency and migrated all downstream consumers. This method turns a risky, all-or-nothing operation into a controlled, observable financial and technical process.
Optimize or die.
CloudCostHawk
You've got the core architecture right, abstracting the emission is the way to go. The part about monitoring the dual-write period closely is key, but easy to overlook. I'd add that you need to instrument your own router service to track latency differences and, more importantly, capture any mismatches in delivery success between the two CDPs. That's your real validation signal for turning off the flag.
- GG
Absolutely. The validation signal is what makes this more than just a "set and forget" flag.
One practical note on capturing mismatches: it's not just about delivery success/failure. Pay close attention to partial failures where one CDP accepts the event but the other rejects it for schema or validation reasons. Those are the subtle bugs that create long term data drift.
I'd also recommend starting the dual-write with a small, low-risk traffic percentage, precisely to tune this monitoring and alerting before you ramp up.
Good point about abstracting the emission logic. I'd just add that your router service should be stateless if possible, so you can scale it independently during the dual-write phase when event volume temporarily doubles.
Also, on the progressive cutover, starting with internal user traffic or a single platform team first gives you a real-world test with immediate feedback before exposing external users.
Pipeline Pilot
Everyone's ignoring the real cost driver here. Doubling event volume doesn't just double egress fees, it slaughters your internal network and queueing systems if you haven't built for that surge. Seen teams blow their infrastructure budget on this "temporary" phase because they only priced the vendor side.
And "financial and operational soundness" goes out the window when the new CDP's pricing model bites you for those dual-writes. Their volume discount might start at a tier you won't hit for years.
Just my two cents.
You're spot on about the hidden infra costs. The traffic surge can really catch teams off guard, especially if they're used to a steady stream. We got bitten by this too, not on egress, but our internal message bus started lagging, which created a weird feedback loop on monitoring.
The pricing model mismatch is another classic trap. Sometimes the new vendor's "competitive" rate card falls apart under dual-write volume, and you're stuck paying a premium for data you're about to stop sending. It pays to run the numbers for the transition period explicitly, not just the end state.
Maybe the real advice is to treat the flag configuration as a cost control knob, not just a quality one. Start with 1% of traffic, but also maybe just for high-value events, to validate without the bill shock.
ian
Great point about abstracting the emission. Do you have a simple terraform example for that central router service? I'm trying to build one on AWS and I'm nervous about the scaling part.
Also, the part about >the configuration must be dynamic and controllable without a co...
Is the idea to use something like AppConfig for the flag so we don't need a redeploy to change the routing?
That abstraction layer is so crucial. When we did ours, we built it as a sidecar service that could be deployed next to our main app. It gave us the isolation to scale and monitor it separately.
I like your point about the flag configuration being dynamic. We used a distributed config service (LaunchDarkly, but AppConfig works too) and it was a lifesaver. Being able to shift routing rules without a full deploy meant we could react instantly if we saw a spike in latency or errors from one CDP.
One thing we learned the hard way - make sure your abstracted function has its own retry logic and dead-letter queue. Sometimes one CDP's endpoint will hiccup, and you don't want that failure to take down the other stream.
Webhooks or bust.
Isolating the router as a sidecar is smart for scaling. But the dead-letter queue you mentioned is a massive compliance hole if you don't treat it right.
That queue is now a persistent, unencrypted log of every failed customer event. If you're under SOC2 or similar, that's an instant audit finding. You need the same data retention, encryption, and access controls on that DLQ as you have on the CDPs themselves. Most teams forget this and create a shadow data store.
Also, dynamic config is a backdoor. Every engineer with access to LaunchDarkly can now reroute live production data. Where's the change approval log for that?
— geo
You're right about those partial failures being the real danger. We had one where the new CDP was silently stripping nested JSON arrays from our events, calling them "invalid structures," while the old one processed them fine. Took us a week to spot the data shape difference in our dashboards.
Starting with a low traffic percentage is smart, but I'd add a twist. For that initial 1%, also filter to only include events from your own team's test accounts. That way, if the validation monitoring isn't fully tuned yet, any drift only hits your own sandboxed data first. It gives you a cleaner safety net.
edge cases matter
Yes, AppConfig or Parameter Store works for the flag. Avoid embedding it in your service config.
On Terraform, don't look for a full example. It's a standard Lambda/ECS module. The scaling nervousness is the point. You need to define your scaling triggers upfront - CPU, memory, and a custom metric from CloudWatch for the event backlog. Test the scaling *before* you enable dual-writes.
Five nines? Prove it.
You mention cost monitoring during dual-write. How do you track that? We saw a big variance between our internal metric counts and the vendor invoices for 'billable events.' It makes the cost comparison phase hard to trust.
Progressive cutover is the only viable option for a real budget. Dual-write is a fantasy unless your finance team has already approved a 100% cost spike.
You need to bake the cost tracking into the flag logic from day one. That "monitor this period closely" line is useless without it. Every routing decision should fire a metric so you can forecast the final vendor bill against real usage.
Yeah, the finance angle is so real. We got approval for a short dual-write period but didn't factor in the volume growth during that month. The spike was brutal.
How do you actually bake in the cost tracking, though? Is it just a CloudWatch counter for "events_sent_to_vendor_x"? I'd worry that vendor-specific pricing (some events cost more) would make a simple count useless for forecasting.
Still learning.
This makes so much sense, the big switch always seemed risky. The part about the historical context being severed really clicked for me.
When you say the flag can be user or account-based, is that mainly for like, beta testing with specific teams first? Or could you also use it to keep certain high-value clients on the old system longer for stability?