Alright, let's talk about Fleet's certificate renewal. Or rather, let's talk about how it doesn't talk to you until it's far too late.
Our setup is bog-standard: self-managed Elastic stack, Fleet Server on dedicated VMs, agents across a few hundred nodes. TLS certs for Fleet Server were generated via `elasticsearch-certutil` at install. The documentation, when you can find the relevant scattered pages, implies renewal is handled. It is not. We discovered this the hard way when a subset of agents began failing with opaque TLS errors. The Fleet Server logs? Equally useless. No pre-expiry warnings in Kibana, no prominent alerts, nothing.
The root cause was the classic PKI surprise: the auto-generated certificates had a one-year validity. The renewal mechanism appears to be:
* Vaguely mentioned as "automated" in some contexts.
* Actually requires a manual process involving `certutil` again, restarting Fleet Server, and then hoping the agents pick up the new cert without requiring a re-enroll.
Here's the kicker—the process for renewal isn't even documented alongside the main Fleet Security setup. You have to go digging. For an infrastructure component critical to security visibility, this is a glaring oversight.
Has anyone else been blindsided by this? What's your operational procedure to avoid the inevitable midnight expiry?
- Do you have a monitoring check on the `fleet.crt` expiry?
- Have you scripted a renewal process that doesn't require agent redeployment?
- Is there *any* native alerting within Stack Management we all missed?
The silence from the logs before the failure is the most concerning part. For a platform built on observability, this feels like a significant unobserved failure mode.
- Nina
- Nina