Structuring a martech stack for a single brand is a well-documented challenge, but introducing multiple brands—each with potentially distinct audiences, regional focuses, and business models—exponentially increases complexity. The core tension lies between enforcing governance/consolidating cost and allowing brand autonomy for agility. Based on my work implementing these systems for a portfolio of 7 B2C and B2B brands, the architecture must be built on a centralized data foundation with federated execution layers.
The guiding principle is **centralized identity, centralized data, federated activation**. The stack should be inverted from the traditional brand-centric tool silo model.
**Core Centralized Layer (Shared Services):**
* **Data Warehouse/Lakehouse:** This is the non-negotiable heart. All customer touchpoint data from every brand must land here. I strongly recommend a single cloud data platform (BigQuery or Snowflake) for unified compute and storage governance.
* **Identity Resolution & CDP:** A single instance stitching user identities across brands is critical for understanding cross-brand journeys. This can be a dedicated tool (e.g., Segment, mParticle) or a homegrown solution on your data platform. The golden record lives here.
* **Master Data Management:** A centralized repository for critical shared dimensions like product catalogs (if applicable), standardized business definitions, and taxonomy.
* **Data Pipeline Orchestration:** A single orchestrator (e.g., Dagster, Airflow) managing data flows from all source systems into the centralized warehouse.
**Federated Activation Layer (Brand-Specific):**
* **Analytics & BI:** A centralized BI platform (e.g., Looker, Mode) with separate workspaces and data models per brand, sourced from the shared warehouse.
* **Marketing Execution:** Here, autonomy is key. Each brand may use its own ESP (e.g., Braze for Brand A, Klaviyo for Brand B), ad platform accounts, and CMS. The critical rule is that all *outbound* event data from these tools must be captured back into the central warehouse via tracking pixels or API exports.
**Technical Implementation Pattern:**
The data model uses a composite key of `(brand_id, user_id)` across all tables. Views or materialized tables are then built per brand as secure views in the warehouse, providing each team a filtered dataset.
```sql
-- Example of a centralized events table and a brand-specific view
CREATE TABLE `centralized_events` (
event_timestamp TIMESTAMP,
brand_id STRING,
user_id STRING,
event_name STRING,
event_properties JSON
);
CREATE VIEW `brand_alpha_events` AS
SELECT * EXCEPT(brand_id)
FROM `centralized_events`
WHERE brand_id = 'alpha'
-- Add row-level security or use separate datasets for production;
```
**Recommendations by Business Context:**
* **High-Volume, Transactional Brands (1M+ MAU):** The cost argument for centralized infrastructure is overwhelming. Use a single, powerful warehouse. Consider a single, enterprise ESP with strong segmentation capabilities to manage costs, even if send-time logic differs by brand.
* **Low-Volume, High-Touch B2B Brands:** The focus shifts to lead routing and account-based marketing. A centralized warehouse is still paramount, but you may tolerate more tool duplication (e.g., different marketing automation platforms) if integration costs are lower than retraining teams.
The primary failure mode is allowing brands to choose their own analytics or CDP tools, which fragments the customer view and eliminates cross-brand learnings. The stack's success is 30% technology and 70% governance: establishing clear data contracts from brand tools to the central platform and managing the change management with brand teams.
--DC
data is the product
You had me until "Identity Resolution & CDP." Now we're in hype territory.
That centralized data warehouse is the real CDP. Just write the entity resolution logic in SQL/DBT, don't buy a separate tool that'll just add another vendor, another point of failure, and another bill. It's just another layer to transform your clean data into their proprietary schema before you pull it back out again.
I've seen this movie. The "single instance stitching" tool becomes a bottleneck. Let the warehouse do the stitching.
SQL is enough
"Non-negotiable heart" is a bit much. A single central warehouse becomes a single point of failure and a political bottleneck for any brand that needs to move fast. You try telling the hotshot new brand team they can't ship a campaign because their schema change is stuck in your central governance queue.
The real trick isn't just building the monolith, it's enforcing that every brand actually sends *all* their data into it. Good luck getting finance to share their lead list.
-- old school
You're right about the enforcement being the real battle, but you're missing where the bottleneck actually forms.
It's not the warehouse, it's the access controls and the auditing. If a brand team can't get their schema change approved quickly, your governance process is broken, not your data architecture.
The failure I've seen isn't a central monolith, it's brands quietly signing up for their own analytics SaaS because "it's just for a pilot." Now you've got data completely off-book, with worse security. A single warehouse can still have decentralized pipelines - the key is treating it like a platform, not a cathedral.
That's a crucial distinction you're making. "Treating it like a platform, not a cathedral" is exactly right. The governance process is the service level agreement, not a doctrine. If a brand team's request gets stuck in a two-week queue, they will go around it, and that's a failure of the platform team to provide adequate service.
The political challenge you identified, with brands spinning up off-book SaaS for a pilot, is often the direct result of that platform service being too slow. The cost of those shadow tools isn't just the license fee, it's the complete loss of data lineage and the security risk you mentioned.
Keep it civil, keep it real
That SLA mindset is the only way to scale governance without rebellion. I've pushed for setting explicit, published service tiers for the central platform.
For example:
- **Tier 1 (Critical):** New event stream for checkout funnel. SLA: 3 business days.
- **Tier:(Standard):** New property on an existing event. SLA: Same day.
- **Tier 3 (Exploratory):** New data source for a 6-week pilot. SLA: 5 days, but with a sunset clause.
If you miss the SLA, the brand team gets a ticket to go use an approved external tool for that specific need. It turns a political fight into a service failure metric for the platform team. You're right, the shadow IT isn't malice, it's a symptom of a platform that's too slow.
Latency is the enemy, but consistency is the goal.
I love this starting point - **centralized identity, centralized data, federated activation** is such a clean mantra to build from. But where I've seen this model get messy is right at the start of that second sentence: "All customer touchpoint data from every brand *must* land here."
The friction point becomes those sneaky, high-value first-party datasets that aren't technically "customer touchpoints" in the classic martech sense. Things like in-app telemetry from a proprietary software product, or detailed fulfillment data from a logistics brand. The product team sees it as *their* operational data, not marketing's data. They'll fight tooth and nail to keep it in a separate "product analytics" silo, even though it holds the key to understanding user behavior before they ever hit a marketing channel.
So my addition would be: that non-negotiable warehouse has to be positioned and *funded* as a *company* asset for all customer-related data, not just a *marketing* asset for campaigns. Otherwise, you're only solving half the cross-brand journey puzzle.
If it's not measurable, it's not marketing.
> "That centralized data warehouse is the real CDP."
I really like this point. It cuts through a lot of the vendor noise.
But what about when you need to push that unified profile *out* to execution tools in real-time? Like syncing a stitched identity to a live ad platform or email tool. Is that still doable just from the warehouse, or do you hit a wall there?
The push for a single cloud data platform like BigQuery or Snowflake is the correct starting point, but you haven't addressed the cost dimension of that "unified compute and storage governance." A single platform can consolidate costs, but without explicit financial architecture, it becomes a massive, opaque shared bill leading to internal chargeback disputes.
You must implement a tagging and allocation strategy from day one, using labels or tags for every dataset, pipeline, and query that identify the owning brand and cost center. This allows you to report on consumption by brand, which is critical for accountability. I've seen platforms where the central data team becomes a cost center villain because brands can't see their own usage. The governance isn't just about data access, it's about financial transparency.
Every dollar counts.
Your friction point is real, but the funding argument is where it falls apart. Positioning it as a "company asset" just means the cost gets buried in a corporate IT budget that nobody owns.
The product team fights to keep their data separate because they don't want to pay for the compute. The moment you try to charge them for storage and queries on that telemetry data, you'll get the real objection. "Why should our P&L fund the marketing team's segmentation jobs?"
The only way this works is if the central platform has a clear, automated chargeback model from day one. Show each team their own line item. Otherwise, it's just a political fight disguised as a data strategy.
show me the bill
Chargeback models are a fantasy if you can't define a "unit" that teams accept. What's a fair cost for a telemetry event? The marketing team running giant joins on it or the product team doing a simple count?
You'll spend more engineering hours building the metering and fighting over the invoice allocations than you ever saved with centralization. Seen it blow up twice. The teams that care about cost just go back to their own Postgres instance where the bill is predictable.
Your vendor is not your friend.
You're hitting the nail on the head about the metering problem. The chargeback fight is real, but the alternative - a big, opaque bill - is what kills platform adoption for good.
I've had some luck treating the central platform's cost as a fixed infrastructure tax, like AWS networking. Then, you only do chargebacks for *extraordinary* consumption, like a brand's runaway query that spikes the monthly bill by 30%. It stops the daily squabbles over pennies and focuses the conversation on actual waste.
That said, teams will still go back to their own Postgres if the central platform is slow or unreliable. Cost visibility is secondary to getting their work done.
Prompt engineering is the new debugging