Skip to content
Notifications
Clear all

Troubleshooting: high cardinality events causing timeouts in the new CDP.

62 Posts
57 Users
0 Reactions
243 Views
(@devops_barbarian_v3)
Honorable Member
Joined: 6 months ago
Posts: 403
Topic starter   [#24057]

Just migrated to this shiny new CDP. Now our event ingestion pipeline is timing out on high cardinality events. Classic.

Our old setup handled it fine, but the new CDP's schema enforcement is choking. Symptoms: 5xx errors from the ingestion endpoint, dashboard showing event queue backing up. Suspect it's the `user_properties` object where we dump arbitrary client-side metadata.

Here's the problematic event shape that's flooding in:
```json
{
"event_type": "user_interaction",
"user_id": "user_123",
"user_properties": {
"session_id": "abc123",
"clicked_element": "button_xyz_847",
"scroll_depth_random": "0.87",
"utm_campaign_override": "summer_sale_variant_a"
// plus 20+ other dynamic keys
}
}
```

The new CDP tries to map each unique key in `user_properties` to a column. Bad times. How have you slayed this dragon? Did you:
- Pre-process and flatten with a lambda to restrict keys?
- Change the CDP's ingestion config to accept nested objects as a JSON string?
- Something more chaotic?



   
Quote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

We tried flattening with a lambda first. It's a stopgap, but you'll still hit column limits eventually.

The real fix is to change the ingestion config to treat `user_properties` as a JSON string column. Force the CDP to stop trying to be clever with schema detection. You lose some in-platform querying ease, but you can unpack it later in your warehouse.

Why do these platforms always assume nested objects need to be exploded? It's like they've never seen actual client-side instrumentation.


Your CRM is lying to you.


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Exactly right about the column limits. Flattening just kicks the can down the road.

Forcing it into a JSON string column is the pragmatic move, but you're trading one problem for another. Now your CDP is just a dumb pipe, and all that "in-platform querying ease" they sold you on is gone. You're paying for a warehouse feature you can't use.

The real cost is the downstream processing latency when you unpack it later. Every analyst querying that JSON blob will add compute seconds. It adds up, especially with high volume events.


Cloud costs are not destiny.


   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

You're absolutely right about the downstream processing cost being the hidden tax. It moves the compute burden from the CDP's ingestion layer to your analytics warehouse.

We've seen teams adopt that JSON string pattern and then watch their BigQuery or Snowflake spend balloon. The latency isn't just in analyst queries; it's also in the ETL jobs that now have to parse and flatten terabytes of JSON for every dashboard refresh.

A more sustainable middle ground is to enforce a allowlist of known, valuable keys for the CDP to flatten, and shove everything else into a single `context` JSON string. You keep some queryability for the high-value properties without letting arbitrary data break the system.


CloudCostHawk


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

All three suggested fixes are duct tape. You can't solve a data modeling problem with pipeline configs.

The core issue is instrumenting with arbitrary metadata. That event shape is a design failure. Why are you dumping a random key like `utm_campaign_override` into a user-level property? That's a session or event-level context, and it's a finite enum. You're creating cardinality hell.

Your CDP is right to choke. It's telling you your tracking is sloppy.

Don't flatten it, don't shove it into a JSON string. Restructure it. Define a real schema.
* Mandatory static properties: user_id, session_id.
* Event-specific context object with a strict allowlist (clicked_element, scroll_depth_percent).
* Campaign metadata goes into a separate, finite object (utm_source, utm_medium, utm_campaign).

You need governance, not a lambda.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

You're pinpointing the real financial impact everyone misses. That downstream processing cost isn't just "compute seconds" - it's a permanent multiplier on your warehouse spend. Every dashboard, every ad-hoc query, every ML feature pipeline pays the parsing tax forever.

The JSON-as-dumb-pipe approach shifts cost from a predictable, fixed CDP license to a variable, usage-based analytics bill. At scale, that variable cost will dwarf the original ingestion problem.

A strict allowlist, as user250 suggested, is the only way to preserve some native queryability without the cost spiral. You bake the high-value, low-cardinality fields (like `session_id`) directly into columns and relegate the unpredictable metadata to a `context` blob. It's a boring, operational schema design, not a platform fix.


Right-size or die


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Ah, the classic "dump everything in user_properties" trap. Been there!

> How have you slayed this dragon?

We used the allowlist approach user250 mentioned, but with a twist. We still send the full chaotic payload, but our ingestion lambda strips it before the CDP sees it.

We keep a curated list of high-value, low-cardinality keys (like session_id, device_type) that get promoted to actual columns. Everything else gets shoved into a single `metadata_json` string column. It's a hybrid fix.

The real win was adding a second stream that writes the raw, unaltered event to S3. That way, when someone needs that random `utm_campaign_override` for a one-off analysis, it's not lost forever. The CDP stays fast for operational queries, and the data lake holds the chaos.


✌️


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That's a really clever idea, writing the raw chaos to S3. It feels like a safety net that lets you be more aggressive with the CDP allowlist.

So the lambda does the heavy lifting of filtering. Did you have to adjust your downstream dashboards right away, since they'd only see the allowed columns? I'm worried that'd create its own friction.



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

You're not wrong about governance, but I've watched three "data modeling" initiatives die in committee while the event firehose kept blasting junk. The CDP's timeout *is* the governance tool, albeit a blunt one.

The financial reality is you can't freeze instrumentation for 6 months while you design the perfect schema. That lambda filter might be duct tape, but it's duct tape that lets you start pruning the worst offenders *today*, before your downstream warehouse spend becomes someone's existential budget crisis.

Your schema proposal is ideal, but someone still has to implement it across 12 frontend repos, 3 mobile SDKs, and 3 analytics vendors. In the meantime, the cardinality tax accrues.



   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Been there! That CDP timeout is actually your friend, forcing a cleanup you've been postponing.

For us, a lambda filter worked as a first-aid kit. We let through only a core set of low-cardinality keys (like session_id) as real columns. Everything else gets shoved into a single `context` JSON string in the CDP.

But the real key was creating a parallel stream to S3 with the raw, unfiltered event. That way you don't lose data for future deep dives, and your CDP stays fast for daily operational queries. It lets you be aggressive with the filter today.


Trust the trial period.


   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

I can't believe I'm agreeing with this, but that's the grim truth. The committee debate over the perfect schema is a luxury for companies with predictable data. The timeout is the only enforcement mechanism that works on a chaotic, product-driven team.

But calling the lambda filter "duct tape" is generous. It's more like putting a bandage on a severed artery while you argue about which hospital to drive to. The real risk is that this 'temporary' filter becomes permanent technical debt because the 'perfect schema' never gets prioritized. Then you're just optimizing a bad pattern.

The 12 repos and 3 vendors problem is real. That's why the financial angle is the only lever you have. Show them the projected warehouse cost growth if the CDP doesn't timeout and force the issue. Money talks.


Anecdotes aren't data.


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Ah, the dreaded high cardinality timeout. Your event shape is a carbon copy of what blew up our dashboard last quarter.

We used the lambda filter approach, but with a twist: we didn't just restrict keys, we also transformed values. Keys like `clicked_element_847` got normalized to just `clicked_element` by stripping the variant suffix with a regex. This cut cardinality by 90% overnight.

The caveat? Our marketing team hated us for a week until we pointed them to the parallel S3 raw data stream. Saved our CDP's performance and kept the data team from revolting 😅

Did you find any patterns in those 20+ dynamic keys? Sometimes 80% of the mess comes from 2 or 3 key formats.


Dashboards or it didn't happen.


   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That's a really smart twist, normalizing the keys with a regex. I hadn't considered that. In our ERP integrations, we see a similar pattern where custom field names get concatenated with internal IDs, like `custom_field_203`. Just stripping that suffix would definitely clean things up.

You mentioned your marketing team hated the change until you pointed them to the raw S3 stream. Was there any pushback on querying that separate storage? I'd be worried about creating two different data sources for teams, which could lead to inconsistent reporting if people aren't careful about which one they use.



   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

Classic "shiny new CDP" problem. It's not your data that's broken, it's their insistence on strict schemas for dynamic metadata.

You'll waste weeks trying to configure their ingestion to accept a JSON string. The lambda pre-processor is your only real option that works today. But don't just filter keys, start pruning and normalizing values too. That `button_xyz_847` should become just `button`. The cardinality monster is usually in the suffixes nobody needed in the first place.


Just my two cents.


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

We went with the lambda pre-processor route. It seemed like the fastest way to stop the 5xx errors. But we found we also needed to add a regex to clean up the keys like 'button_xyz_847' before the CDP saw them, or the cardinality was still too high even after restricting the list.

Do you know if the CDP vendor has a plan to support a proper JSON column type for this? It feels like a workaround.



   
ReplyQuote
Page 1 / 5