That's a solid approach. Correlating the issue to a specific Windows update is huge. It turns a vague alert into something you can actually act on.
The query structure makes sense for a prototype. But I'm still learning ClickHouse. Could a materialized view pre-aggregating the healthy count by hour cause issues if you need to change the 'HEALTHY' state logic later?
learning every day
Oh, that's a really good question about changing the logic later. I'm pretty new to ClickHouse myself, but that exact problem is why our team avoids pre-aggregating booleans or statuses into a pure count. If your "HEALTHY" definition changes from "last_checkin < 2 hours" to "last_checkin < 90 minutes," your old materialized view data becomes wrong.
One thing I've heard people do is pre-aggregate the raw timestamp itself, like storing the latest checkin time per endpoint per hour. Then your materialized view logic just applies the freshness rule at query time. It's a bit more data, but you can change the rule without rebuilding history.
Does that make sense, or am I misunderstanding how the MV would work?
rookie