They told us to pick the best tool for the job. So we did. Then the real work began.
The schema mapping and historical backfill were a known quantity. You write the scripts, you test the idempotency, you monitor the queues. The pain was in the 47 "approved" downstream systems that had hardcoded API keys and expectations from the old vendor. Each team owner acted like we were personally migrating their child to a different school.
The compliance angle was our only leverage. Old CDP was storing PII in logs with a 180-day retention "for debugging." New contract required SOC2 Type II and a real data privacy framework. Suddenly, "but our custom integration breaks" met "but this is a compliance finding."
Example: getting the marketing team to accept that their beloved "lead score" field needed to be hashed before leaving our new pipeline. Their connector broke. Good.
```python
# What they had
payload["email"] = user.email
payload["lead_score"] = user.inferred_wealth_score # Seriously.
# What they got
payload["email_sha256"] = hash_for_join(user.email)
# lead_score died a quiet death. No more inferred wealth.
```
The post-mortem has one line in bold: **Define "data ownership" before you define the architecture.** Otherwise, every department with a Python script becomes a stakeholder with veto power. The tech migrates in months. The politics allocate the years.
Trust but verify – and audit