Skip to content
Notifications
Clear all

Just built a simple migration monitor that compares counts hourly - sharing the code

1 Posts
1 Users
0 Reactions
29 Views
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
Topic starter   [#7820]

Just finished migrating our main analytics pipeline from one CDP to another. The actual data transfer went smoothly, but what kept me up was verifying that nothing got dropped in transit—especially with historical events.

I built a simple monitor that compares hourly event counts between the old and new systems. It runs as a Kubernetes CronJob, queries both sources, and pushes diffs to Prometheus so we can alert on significant discrepancies. Here's the core of it:

```python
def compare_hourly_counts(old_cdp_query, new_cdp_query, hour):
# Query both systems for the given hour
old_count = query_cdp(old_cdp_query, hour)
new_count = query_cdp(new_cdp_query, hour)

diff = abs(old_count - new_count)
diff_percent = (diff / old_count * 100) if old_count > 0 else 0

# Push to Prometheus
push_to_gateway('cdp_migration_diff', {
'diff_absolute': diff,
'diff_percent': diff_percent,
'old_count': old_count,
'new_count': new_count
}, hour)
```

Key takeaways:
* The monitor logs full mismatches (zero events in new system) as critical alerts
* Tolerates 5%
* We backfilled the comparison for the last 30 days to catch silent data loss

The dashboard shows three main panels:
- Absolute count difference per hour
- Percentage drift
- Cumulative gap over time

Surprisingly, it caught a few hours where our new CDP's deduplication was more aggressive than expected. Now I can sleep a bit better knowing the pipeline is verified.

Has anyone else built similar validation during CDP migrations? Curious about other methods for ensuring event consistency beyond simple counts.

- away



   
Quote