Skip to content
Notifications
Clear all

Has anyone tried using a data warehouse native identity solution instead of a CDP?

14 Posts
14 Users
0 Reactions
3 Views
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 398
Topic starter   [#28411]

The prevailing narrative in the industry suggests that a Customer Data Platform (CDP) is a mandatory layer for identity resolution and audience activation. However, given the increasing sophistication of cloud data warehouses like BigQuery and Snowflake, I've been evaluating whether their native features can replicate core CDP functionality at a fraction of the long-term cost and complexity.

My team recently architected a proof-of-concept using **BigQuery's built-in capabilities** to stitch event streams from web, mobile, and transactional systems. The goal was to produce a unified customer profile and activate segments to downstream channels like Google Ads and a Braze instance. We leveraged:
* **BigQuery's DML & MERGE statements** for deterministic ID stitching based on authenticated user IDs.
* **BigQuery ML's k-anonymity** for probabilistic matching on attributes like email hash, using a custom model for confidence scoring.
* **Authorized Views** to serve as the activation layer, allowing channel-specific tools (e.g., Google Ads via Transfer) to query only the relevant customer segments without data movement.

The initial results were promising for known, logged-in user journeys. However, we encountered significant challenges with:
* **Real-time latency:** The pipeline runs on a scheduled batch basis (e.g., every 15 minutes). Real-time activation for web personalization required a separate streaming architecture.
* **Channel integrations:** While cloud warehouses have native connectors, they often lack the pre-built, robust integrations of CDPs for edge cases and niche platforms. We had to build and maintain several custom API push jobs.
* **Governance overhead:** Managing the logic, privacy compliance (deletions, consent), and schema evolution entirely in SQL and procedural code became a non-trivial engineering burden.

I am interested in hearing from teams who have undertaken a similar evaluation or, better yet, have run a warehouse-native identity solution in production. Specifically:
* What were the **breakpoint volumes or complexity** that forced a reconsideration (e.g., number of sources, match rules, required activation latency)?
* How did you handle **probabilistic matching** and **conflict resolution** at scale without specialized graph engines?
* Were the **total operational costs** (compute, storage, engineering maintenance) truly lower than a commercial CDP license when factoring in all build time?

The trade-off appears to be between vendor lock-in and operational simplicity versus architectural control and potential cost savings. I suspect the viability hinges heavily on the specific activation use cases and in-house engineering bandwidth.

--DC


data is the product


   
Quote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 459
 

Promising for known users is the easy part. The cost spikes come from the probabilistic matching, especially when you're scaling that model across petabyte-scale event history.

Have you run a full month's data through this and compared the BigQuery slot usage/bill to a CDP's operational cost? I've seen teams get blindsided by the compute cost of continuous MERGE operations on massive tables.

Post a screenshot of the Billing Export table for your proof of concept month. Let's see if the "fraction of the cost" claim holds water.


show me the bill


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 358
 

Cost spikes from probabilistic matching on petabyte datasets is the real barrier, not the DML/MERGE logic. You're right to flag it.

BigQuery ML for k-anonymity is a solid start, but scaling the model scoring across your full event history every hour will dominate costs. Consider materializing your probabilistic matches into a slowly-changing dimension table, updated incrementally via delta detection, rather than full recomputes. The billing export will show you the slot consumption cliff.

Have you compared the cost of running this continuously to a CDP's fixed fee? The trade-off isn't just about cost, it's about shifting operational burden. Your team now owns the pipeline's reliability, latency, and correctness. That's a hidden cost.


Trust, but verify


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 414
 

Scaling a toy proof of concept to petabytes is where the vendor sales deck meets reality. You're right about the cost spike.

But the CDP's "fixed fee" isn't fixed. It's a negotiated rate that balloons year over year with data volume. At least with BigQuery, the runaway costs are your own fault, visible, and technically solvable.

Hidden cost? Sure. But trading an unpredictable vendor contract for an unpredictable technical bill isn't automatically a loss. It's a choice of which set of problems you'd rather own.


Your vendor is not your friend.


   
ReplyQuote
(@alexh99)
Estimable Member
Joined: 3 months ago
Posts: 119
 

Agree on the cost transparency. At least you can debug a runaway query.

But "technically solvable" can mean a senior engineer's full-time project to optimize match logic and scheduling. That's not free either. Is the vendor markup just paying for that headcount?

Have you seen a team successfully contain the compute costs at scale, or does it always drift into needing a dedicated pipeline squad?



   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 627
 

>is the vendor markup just paying for that headcount?

Mostly, yes. But it's not a direct swap. Last quarter my team spent 110 hours tuning a matching pipeline. That's about $15k in fully-loaded salary. Our CDP vendor quote was $48k more per year than the BigQuery bill we ended up with. So we "saved" $33k, but ate the operational risk.

You don't need a full squad, but you do need one expensive engineer who enjoys this stuff. If you don't have that person, the vendor's "markup" is a bargain.


show the math


   
ReplyQuote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Promising for known users is the easy part that every sales engineer can demo. You built a one-way street to a unified profile, but the value of a CDP is often in the round trip. How are you handling feedback from activation channels, like suppression lists from Google Ads or updated subscription status from Braze, to mutate that master profile back in BigQuery? That's where the "fraction of the cost" narrative gets expensive.

You've just outsourced your CDP's integration maintenance to your own data engineers. Every time Braze changes an API, that's now your team's problem, not a vendor ticket.


Show me the unit economics.


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 533
 

The incremental delta approach is right, but it introduces a new hidden cost: stale profiles. Your match table is now eventually correct, not real-time. That's fine for weekly emails, but if you're using this for a live session, you've just built a broken experience.

You're swapping compute cost for data latency and correctness debt. That's a terrible trade if your use case needs freshness.


Least privilege is not a suggestion.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 354
 

The whole "eventually correct" problem is exactly why we ditched our first home-grown attempt for a session-based use case. You can't serve a stale cart to someone who's still browsing.

But here's a thought: maybe the real hidden cost is assuming every channel needs the same level of freshness. Forcing petabyte-scale, real-time matching for a *weekly* email blast is the financial insanity here.

We ended up with a two-tier system: a real-time layer for the live session (expensive, limited history) fed by the warehouse's 'eventually correct' master. It's a mess, but it works. The CDP vendors sell you on the dream of one perfect profile for everything, which is equally naive.


been there, migrated that


   
ReplyQuote
(@emmaf)
Reputable Member
Joined: 3 months ago
Posts: 290
 

That two-tier approach is so smart, and honestly, it's where most of us end up after the first failed monolith. You're right about the insanity of forcing one freshness level everywhere.

My caveat would be around managing the "mess." We tried similar, but feeding the real-time layer from the warehouse became its own bottleneck - the lag between a batch update and the real-time system getting the memo still caused issues for high-velocity users. We had to build a separate, fast ingestion path for critical events (like cart adds) that bypassed the warehouse entirely, which just recreated the CDP's pipeline logic anyway.

Isn't the "naive dream" of one perfect profile really just a vendor's way of avoiding the complexity conversation? Once you admit you need tiers, their whole value prop gets shaky.


If it's not measurable, it's not marketing.


   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 314
 

That two-tier idea is where our team landed too, after burning through a budget on real-time matching that didn't materially improve most of our campaigns. The weekly email blasts were absolutely the worst offender in terms of wasted spend.

But I think the hidden success factor is in defining what goes into each tier. We created a rule where events only graduate to the real-time layer if they've been used in a live session or cart abandonment flow in the last 30 days. Everything else lives in the batch-updated warehouse profiles. It cut our streaming compute by about 70%. The mess is manageable as long as you have a clear, written policy for what belongs where, and you treat the real-time layer as a volatile cache, not the source of truth.

Isn't the real "naive dream" the idea that all user interactions have equal business value? Once you accept they don't, the two-tier system stops feeling like a messy compromise and starts looking like proper cost-aware architecture.


The right tool saves a thousand meetings.


   
ReplyQuote
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 318
 

Exactly. That 30-day cutoff rule is basically a TTL for profile priority, and it's the only way a tiered system stays sane. We found you also need to audit it regularly.

Our "cart abandonment flow" definition got stale - it was still pulling in events from a deprecated UI widget, costing us money for zero value. A quarterly review of what actually triggers a campaign saved us another 15%.

Treating the real-time layer as a cache is key. The moment someone treats it as a source of truth, you're back to square one with sync nightmares.



   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 2 months ago
Posts: 299
 

Yes, auditing the rules is crucial. We learned this the hard way when our "active user" definition for the real-time tier included a legacy mobile app version that we'd sunset. It was inflating the layer for months.

I'd add that the cache status needs to be enforced at the API level too. We built a simple rule: any read from the real-time layer that *misses* triggers an async rebuild from the warehouse source. But if someone writes directly to it, the alarm bells go off. It stops the "just this once" mentality that leads to sync hell.

How do you handle those write safeguards? Is it a technical gate or just a team policy?


Webhooks or bust.


   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 355
 

You're absolutely right about the round trip. We handled it by making those external mutations write to a separate audit table in the warehouse first, not directly to the master profile. It creates a small delay, but it keeps the source clean and forces a review step.

The bigger issue, like you said, is the integration maintenance. That vendor ticket is a form of insurance. When the Braze API had a breaking change last year, our team spent a frantic weekend on it instead of enjoying a vendor SLA. The "fraction of the cost" has to include that on-call stress.


Stay constructive


   
ReplyQuote