Skip to content
Notifications
Clear all

Issue: Fleet's TLS certificate renewal process is opaque and failed for us.

1 Posts
1 Users
0 Reactions
29 Views
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
Topic starter   [#15846]

Alright, let's talk about Fleet's certificate renewal. Or rather, let's talk about how it doesn't talk to you until it's far too late.

Our setup is bog-standard: self-managed Elastic stack, Fleet Server on dedicated VMs, agents across a few hundred nodes. TLS certs for Fleet Server were generated via `elasticsearch-certutil` at install. The documentation, when you can find the relevant scattered pages, implies renewal is handled. It is not. We discovered this the hard way when a subset of agents began failing with opaque TLS errors. The Fleet Server logs? Equally useless. No pre-expiry warnings in Kibana, no prominent alerts, nothing.

The root cause was the classic PKI surprise: the auto-generated certificates had a one-year validity. The renewal mechanism appears to be:
* Vaguely mentioned as "automated" in some contexts.
* Actually requires a manual process involving `certutil` again, restarting Fleet Server, and then hoping the agents pick up the new cert without requiring a re-enroll.

Here's the kicker—the process for renewal isn't even documented alongside the main Fleet Security setup. You have to go digging. For an infrastructure component critical to security visibility, this is a glaring oversight.

Has anyone else been blindsided by this? What's your operational procedure to avoid the inevitable midnight expiry?
- Do you have a monitoring check on the `fleet.crt` expiry?
- Have you scripted a renewal process that doesn't require agent redeployment?
- Is there *any* native alerting within Stack Management we all missed?

The silence from the logs before the failure is the most concerning part. For a platform built on observability, this feels like a significant unobserved failure mode.

- Nina


- Nina


   
Quote