In our recent multi-cloud CDP migration from a legacy homegrown event router to a hybrid Kafka/Google PubSub architecture, we found that operational visibility was the single greatest factor in maintaining velocity and stakeholder confidence. While we had detailed runbooks for schema translation and idempotent backfill processes, the overarching question from leadership was consistently, "How complete is the migration, and what is our risk exposure?" To answer this, we built a comprehensive Grafana dashboard that became our command center.
The dashboard was designed to track three core dimensions: data fidelity, pipeline health, and business completeness. It aggregated metrics from our orchestration layer (Apache Airflow), our streaming platforms, and our data warehouses. Below is a simplified JSON representation of the key dashboard variables we used to segment data by source team, destination cloud provider, and event type.
```json
{
"dashboard_variables": {
"team": ["team_web", "team_mobile", "team_backend"],
"cloud_target": ["gcp_pubsub", "aws_msk", "azure_eventhub"],
"event_schema": ["identify", "pageview", "transaction"],
"timeframe": ["last_24h", "last_7d", "since_migration_start"]
}
}
```
The primary panels included:
* **Event Throughput Comparison:** A dual-axis time-series graph plotting the volume of events per second in the legacy pipeline against the new pipeline, segmented by `team` and `cloud_target`. Discrepancies beyond a 2% threshold triggered a warning annotation.
* **Schema Validation Failure Rate:** A stat panel showing the percentage of events failing Avro schema validation in the new pipeline's ingress point, crucial for catching translation errors early.
* **End-to-End Latency Delta:** The 95th percentile latency difference (new vs. old) for events reaching the data lake. This was our key performance indicator for downstream impact.
* **Backfill Progress:** A cumulative gauge showing the percentage of historical events successfully reprocessed and loaded into the new data model, broken down by `event_schema`.
* **Downstream Connector Health:** A status grid displaying the heartbeats of all re-wired connectors (e.g., Snowflake, BigQuery, Amplitude) with color-coded alerts for any lag or error states.
This dashboard was powered by a telemetry pipeline that emitted custom metrics from our migration workers. The most valuable metric, however, was a derived "Business Completeness Percentage." It was not a simple average. We calculated it as a weighted sum based on the revenue impact of each event type and the completion status of its corresponding team's migration. This gave product leadership a single, risk-adjusted number to monitor.
The implementation forced us to instrument our migration as a first-class observable system, which had an unexpected benefit: the same patterns were retained post-migration for ongoing data quality monitoring. The dashboard evolved from a migration tracker to a permanent data platform health console. I'm interested to hear how others have instrumented complex CDP migrations, particularly when dealing with eventual consistency models across clouds. What metrics proved to be the most leading indicators of a problem?
Boring is beautiful
Nice breakdown. The three dimensions make sense. I'd add a fourth one for cost tracking next time - it's easy to blow the budget during a migration. Did you bake latency from event creation to warehouse arrival into your "data fidelity" metric? That's where we saw the most drift.
Good point on latency. We did include end-to-end latency in the fidelity panel, but we had to break it down by source system. The legacy router batches would cause spikes that masked average health, so we tracked P95 and P99 separately.
Cost tracking is a great addition. We actually had a separate Cloud Cost Management dashboard but didn't think to merge them. For migrations, seeing spend spike in parallel to progress would've been useful context.
Run it yourself.
That variable setup is a good start, but you're going to hit scaling issues if you rely solely on static JSON arrays for your team and schema variables. Once you're past ten teams or start dynamically adding new event types, maintaining that list is a chore.
You need to source these variables from wherever your service catalog or schema registry lives. We use a simple Python exporter that scrapes our internal API and pushes the current values to a Prometheus metric as a gauge with a label. Then you can use the `label_values()` function in Grafana to populate the dropdown dynamically. It looks like this in the dashboard JSON:
```json
"team": {
"query": "label_values(cdp_teams, team)",
"type": "query"
}
```
It saves you from manually updating the dashboard every time a new service onboards.
Automate everything. Twice.
Oh that's clever, sourcing the variable values from a metric! We ended up halfway there by querying our Clickhouse metastore for a live list of event types, but I love the simplicity of a dedicated exporter. It's one less dependency on the warehouse.
My only caveat would be to make sure your scraper is really robust, because if it breaks, your dashboard variables go empty and the whole view collapses. We had a similar thing happen once when our metastore query timed out - suddenly the whole team thought migration had stopped! 😅 A good health check on the exporter/metric itself is a must.
That's a great point about the exporter being robust. How do you handle the health check for something like that? Do you have a separate alert on the metric's freshness, or does your scraper push a separate "last successful scrape" timestamp?
And yeah, a blank dropdown would definitely cause panic! 😅 Maybe having a static fallback list in the variable config would help?
Sourcing dashboard variables from a live system is indeed a more scalable pattern than static JSON arrays. However, the critical dependency it creates warrants a more rigorous monitoring approach than a simple health check.
We've implemented a two-layer verification system for such exporters. First, the exporter itself emits a metric with the timestamp of its last successful scrape, `cdp_variable_scrape_timestamp`. We alert on that metric's staleness using Prometheus's `timestamp()` function. Second, we also alert on the cardinality of the target metric itself, `cdp_teams`. If the number of distinct `team` labels drops below a historical baseline threshold, say a 50% decrease, it signals a potential data failure even if the scraper is technically running.
A static fallback list, as you suggest, is a pragmatic safety net for the dashboard itself. We achieve this by defining the variable with both a dynamic query and a `constant` option in the JSON. If the query returns empty, Grafana can default to the predefined list, preventing a total UI collapse. This configuration requires careful testing of the failure scenario to ensure the fallback triggers correctly.
This is really helpful, seeing a concrete example of the variables you structured the dashboard around. The three dimensions are a solid foundation, and it's clever to use those variables to slice the same underlying metrics for each view.
I'm curious about the decision to include event schema as a variable. In our NetSuite integrations, we track completeness by functional module or transaction type, like orders versus inventory adjustments. Did you find that segmenting by individual event schema became noisy, or was it critical for pinpointing issues in specific data contracts?
The timeframe variable is interesting too. For a long-running migration, did you find yourself mostly using the last 24 hours, or did the longer 7-day window become more useful for spotting trends in backfill progress?
Tracking P95/P99 for latency instead of just average is such a smart move for migrations. We got burned on that once by only watching the average, and it hid a nasty tail latency issue where some events took hours due to queue backlogs.
A cost overlay would have been perfect for context. Did you find that the cost spikes correlated with any specific backfill activity, or was it more from the dual running of old and new systems?
Data is the new oil - but it's usually crude.
You're adding complexity to solve complexity. Now you have to monitor, alert, and maintain this two-layer verification system for a dashboard variable.
The constant fallback is a good trapdoor, but you're still betting your critical migration view on a scraper's health. If your service catalog API is down, your cardinality alert fires. What does the team do? They can't fix the API. They just get an alert while the dashboard is blind.
Sometimes static lists are fine. They're dumb, but they're stable. For a migration dashboard, that's often better than clever.
read the fine print
Glad someone brought up cost, but it's a distraction if you just slap it on the dashboard. You need to tie it to progress, otherwise it's just noise. If cost spikes but the migration velocity hasn't changed, your cloud team gets alerts for no reason.
On latency, tracking it as a fidelity component is the right idea, but I'd be wary of blending it with correctness metrics. You don't want a slow but correct pipeline to tank your overall "health" score and trigger a rollback. Those need separate thresholds.
Question everything
Love that three-dimensional breakdown, it's exactly the kind of structure that gives non-technical stakeholders something concrete to latch onto.
The business completeness dimension is such a good call. In our onboarding platform migration, we set up a similar "percent of active users in new system" metric and it immediately became the CEO's favorite chart. It translates tech progress into business outcomes.
Curious about the "risk exposure" angle - did you bake any specific metrics for that into the dashboard, like failure rate trends or lagging event counts, or was that more of an overall read from combining the three views?
The approach is sound, but your implementation is brittle. Sourcing variables from a live metric just shifts the maintenance from JSON to your scraper's uptime and your API's stability.
That `cdp_teams` gauge will also cause query performance issues as cardinality grows. Every time you open the dashboard, Grafana fires a `label_values()` query that has to scan that entire high-cardinality metric series. It's fine for a few dozen labels, but it's a hidden tax.
Better to expose the list via a dedicated, low-cardinality metric like `info{catalog_team="..."}=1` or, even simpler, have the exporter write a static JSON file to an object store and source the variable with a `jsonList` type. It's still dynamic, but you're not hammering Prometheus for metadata.
—davidr
You're spot on about schema segmentation potentially adding noise. In our case, the event schemas corresponded to very distinct business objects with different validation rules, so having that granularity was essential for isolating contract failures. For a migration focused on transaction types, like your NetSuite example, grouping by functional module is probably the cleaner approach.
Regarding timeframe, the 7-day view was indispensable. A 24-hour window would often show "100% complete" for all schemas, masking the fact that some backfill processes were lagging by days. The longer window surfaced those creeping gaps, especially for low-volume but critical event types.
You're right about the query tax. Using `label_values()` on a high-cardinality gauge is a bad plan. I've seen it grind a dashboard to a halt.
But the static JSON file in an object store just moves the availability problem. Now you're betting on your object store and its permissions. It's the same dependency with extra steps.
Your `info{catalog_team="..."}=1` metric is the real fix. Low cardinality, single scrape, and you can still alert on its presence. We use that pattern for service discovery. It works.