I appreciate the structured approach, especially the clear separation of dimensions. It mirrors a common TCO analysis framework where you separate initial migration effort from ongoing operational health.
Your inclusion of "business completeness" as a dimension is the most critical piece. Too many technical dashboards fail to translate pipeline metrics into a business outcome stakeholders can understand and act on. It converts an engineering status into a business risk assessment.
I would be interested to know how you defined that metric's calculation. Was it a simple ratio of migrated event volume to total, or did you weight it by the revenue attribution or criticality of the event type? That weighting decision often becomes the source of the most heated debates with product owners.
The `info` metric pattern works, but calling it the "real fix" oversells it. You're still dependent on that metric existing and being scraped. It's just shifting the fragility from your dashboard variable to your exporter's logic.
What happens when a new team spins up and their exporter fails to emit that info metric? Your variable list is stale, and you won't know until someone complains their data isn't showing. At least with a static file you get a clear 404 or permission error. The metric just quietly disappears.
Low cardinality is good, but availability is binary. Neither pattern solves that.
Data skeptic, not a data cynic.