I'm planning our migration from Segment to a new CDP and hit a snag. Our source events have deeply nested JSON in the properties (like `product.categories[0].name`). The old platform just flattened it.
My worry is that if I just map the top-level object, the new CDP might treat it as a simple string and I'll lose the ability to query those nested fields later. Has anyone done this before?
What's the best way to preserve this structure during the migration? Should I transform the events before sending them, or does the new CDP usually handle nested JSON natively? I'm using a basic Terraform setup to manage the pipelines.
You're right to be concerned about flattening. Many CDPs that treat nested JSON as a string create a reporting dead end. The native handling question is key.
You need to check the destination's spec. Some platforms, like RudderStack or mParticle, have explicit support for nested objects in their tracking calls. Others might require pre-processing. I'd test sending a sample event with a nested property and then query it in the CDP's interface before scaling the migration.
Given your Terraform setup, you could stage a transformation step in the pipeline if the new CDP doesn't support it natively. But that adds complexity. Start by confirming the destination's capabilities.
Measure twice, spend once
That nested JSON issue is a classic migration pitfall. You're right to avoid stringifying it; querying becomes impossible.
I've seen teams use a middle layer like a Lambda function in the pipeline to reshape data before ingestion. Since you're in Terraform, you could prototype a step that checks the destination's API response for nested fields. If it rejects them, you've got your answer before moving all the traffic.
What's the new CDP? Their documentation on event structure usually gives a definitive yes or no on nested object support.
Good point about checking the API response. A quick test call would save so much time over assuming it works.
I'm curious, when you mention a Lambda function, would you just use it to flatten the data if needed, or try to preserve the nested structure somehow? I'm still learning how to keep that query ability in a new system.
Thanks for sharing this.
>I'm curious, when you mention a Lambda function, would you just use it to flatten the data if needed, or try to preserve the nested structure somehow?
Great question! Honestly, I'd only flatten as a last resort. The whole goal is to keep that structure queryable, right? 😅
In my last migration, I used a Lambda to do a kind of "structure check" on the destination first. If the API accepted nested JSON natively, I'd pass it right through. If it choked, I'd have the Lambda re-shape the data into a format the CDP *could* query properly, like expanding nested properties into separate top-level fields with dot notation. That way you're not just making a string blob.
But it adds pipeline complexity, so definitely do that API test call first to see if you even need it!
Backup first.
That's a really common snag when moving platforms. We hit something similar with nested product data in our analytics pipeline.
You mentioned using Terraform, right? One thing that worked for us was adding a small test stage that sends sample nested events to the new CDP's staging endpoint first. You can bake that check into your Terraform plan as a local-exec provisioner before the full cutover, just to see how the destination API responds. That saved us from a nasty surprise later.
Which CDP are you migrating to? Their docs sometimes have explicit examples for nested structures.
Learning by breaking
That approach with the Lambda is smart, especially the conditional logic. I did something similar but ran into an edge case where the dot notation expansion created duplicate property names from different nested paths. Had to add a small prefix to keep them unique.
Your point about testing the API first is key. I once wasted a week building a transformer only to find out the destination's "staging" endpoint behaved differently than production. A simple script to compare the two API responses would have saved me.
Your Terraform setup is a red herring here, honestly. Infrastructure as code doesn't solve semantic data mapping. I see everyone jumping to Lambda functions and transformation stages, but that's a classic case of solving the wrong problem first.
You need to dig into the destination CDP's actual data model before you write a single line of transformation code. Many of them claim to "accept" nested JSON but then silently stringify it in storage, which is exactly the reporting dead end you're afraid of. The API might return a 200 OK while still wrecking your schema. Check their raw query interface, not just the ingestion spec.
If the new platform can't natively query nested properties, you're better off flattening strategically in your source system now, rather than building a complex pipeline to preserve a structure the destination can't use. Sometimes migrating off a legacy pattern is the real win.
Your k8s cluster is 40% idle.
You've hit on the foundational problem of vendor lock-in disguised as a data pipeline issue. Your comment about Segment flattening the nested JSON is the critical clue; you're inheriting a pre-damaged data model.
The correct sequence isn't building transformation layers yet. It's to immediately validate the new CDP's *query* capabilities, not just its ingestion API. Request a temporary sandbox account and run two tests: first, send a nested event and attempt to query `product.categories[0].name` directly in their SQL or reporting interface. Second, export the raw event back out via their API to see if the nested structure is preserved in storage. Many platforms accept nested JSON on ingress only to serialize it as a string in the data warehouse, which is functionally identical to what you're trying to escape.
If the new platform fails this test, you're facing a product limitation, not an engineering challenge. The strategic decision then shifts to whether you should pre-process the data to align with the new system's flat model, or reconsider your CDP choice altogether. Adding pipeline complexity to fix a destination's weak data model is rarely a good long-term investment.
Good call on RudderStack and mParticle having explicit support. That's the first thing I look for in a spec.
One caveat from a migration I did last year: even if the spec says they support nested objects, you need to verify how they handle *updates* to those nested structures in their identity graph. Some platforms merge them, others overwrite the entire object, which can silently drop data.
Testing with a sample event that includes updates to a nested property, not just a simple send, can expose that difference early.
Ship fast, measure faster.
That's a crucial test case I hadn't considered. The update behavior is often buried deep in their docs, if it's documented at all.
I'd add that you should test a *partial* update to a nested object (e.g., adding a new key to an existing object) versus sending a completely new nested structure. The merge logic can differ between those two scenarios in some platforms.
Great point on checking the identity graph specifically, not just the event stream. That's where data loss really hurts.
—Anita
Terraform won't save you. Your problem started when Segment flattened it. You're trying to fix a broken data model in the middle of a pipeline, which is like repainting a car after the engine fell out.
Everyone's telling you to test the new CDP's API. That's obvious. The real question is why you'd trust another black box platform to handle nested data correctly when the last one didn't. You're just trading one set of unknown behaviors for another.
The "best way" is to stop generating nested JSON you can't control. Flatten it at the source where you own the logic, before any CDP ever sees it. Otherwise you're building a Rube Goldberg machine to appease vendor quirks.
If it ain't broke, don't 'upgrade' it.
> check the destination's spec
That's the step most people skip, and then they're stuck building a transformer for a platform that might've handled it fine. Even when the spec says it's supported, I've seen platforms choke on arrays inside nested objects, or null values in deep properties.
Before I'd touch Terraform, I'd write a one-off script to send a few test events with gnarly nesting - think empty objects, mixed types, arrays of objects - and then immediately try to query them in the UI. If you can't filter a report by `event.properties.product.skus[0].id`, the spec lied.
YMMV
That point about exporting the raw event back out is such a good test. It's the only way to know what's *actually* stored versus what the UI shows you.
I'd add that you should also check how the export API behaves for batched historical data, not just a single event sent in real-time. Sometimes the real-time stream preserves structure, but the bulk export for backfills flattens it, which creates a nasty inconsistency in your data lake.
Raise the signal, lower the noise.
The "best way" is to stop looking for a single best way. You're trying to solve a data modeling problem with pipeline engineering. Terraform doesn't matter here.
The new CDP probably doesn't handle nested JSON natively, or at least not in a queryable way you'd expect. They all claim they do until you try to run a cohort on `properties.product.categories[0].name` and get a type error.
Don't transform before sending. Send a raw nested event to their staging API right now, then immediately try to query it in their SQL editor. If you can't, you have your answer: flatten at source and own the logic. Otherwise you're just building a dependency on their undocumented parsing quirks, which is exactly how you got into this mess with Segment.