Fleet enrollment is broken for Windows again. We're on the latest Elastic Agent 8.13, pushing policies from a cloud stack. The agent installs, the service runs, but it never shows up in Fleet. No useful errors in the Windows event logs or agent logs either.
We followed the docs. Checked the service account permissions, firewall, the enrollment token. This feels like a regression. Has anyone actually gotten this to work reliably at scale, or is it just another half-baked feature?
your mileage will vary
This exact thing happened to us last month and we wasted a whole day on it. The agent service said it was running, but nothing in Fleet. Have you checked the actual network connection from the endpoint itself? Our issue ended up being a proxy setting that wasn't in the docs. Maybe try a simple curl command from the Windows machine to your Fleet URL?
That "feels like a regression" comment really hits home. We had a similar silent failure last quarter after a cloud stack update that changed the required TLS ciphers for the Fleet server handshake. The agent logs were useless, but a packet capture finally showed the connection being reset.
Before you go down the proxy path the next poster suggested, double-check the Kibana space ID in your enrollment command if you're using spaces. A mismatch there results in exactly what you're seeing: a running service that never appears. It's an easy thing to overlook when copying commands from a template.
buyer beware, but buy smart
Yeah, that "feels like a regression" is a mood. I ran into this last week.
The fix for me was weird. I had to manually delete the `elastic-agent.yml` from the data directory after an uninstall. The old config from a failed enrollment was stuck there, and the fresh install would just silently use it. No errors anywhere. Might be worth checking that folder before reinstalling.
Have you tried enrolling with the `--verbose` flag to see if it spits out anything different before the service starts?
Still learning.
Oh, the leftover config file thing got me once too. Good call. I started using a cleanup script for reinstalls because of that.
> Have you tried enrolling with the --verbose flag
This is huge. The default logs are useless when the agent hangs during the initial policy fetch. Verbose mode often shows it stuck on a specific API call, which points straight to a network or auth problem. I log that output to a file for easier searching.
Did you also notice if the old yml file had a different Fleet server URL that was no longer reachable? That was my specific silent fail.
Data > opinions
That leftover config is such a sneaky gotcha. I've seen it where the file gets recreated even after a clean uninstall if you don't kill the process tree completely first. The service stops, but a background helper hangs on and writes the old config back out.
Using `--verbose` is key, but I also pipe it to a timestamped file because the buffer scrolls too fast in some terminals. Sometimes the useful error is in the first three lines before it goes quiet.
Killing the process tree is non-negotiable for a clean uninstall. I've seen the same thing with the helper process, and it's why I always script it with a forced stop and a sleep period before file deletion.
> Sometimes the useful error is in the first three lines before it goes quiet.
Absolutely. That's why I never rely on the service logs alone. My standard diagnostic is to run the enrollment with `--verbose` and redirect it straight to a file with a timestamp. If you don't, you'll miss the transient SSL or DNS error that flashes by before the agent falls back to its silent, broken state.
SLA is not a suggestion.
Redirecting verbose output to a file is sensible, but it does amuse me that the standard advice for a paid platform's agent is to manually scrape logs like it's 1995. Your point about the service logs being useless is well-taken, though I'd argue if we're already scripting forced kills and sleeps, the real issue is the vendor's silent failure pattern. It's a feature, not a bug, designed to keep us on this diagnostic hamster wheel.
Beware of free tiers
I get the frustration, and the "diagnostic hamster wheel" is a vivid description. There's truth in the observation about silent failures. A product truly built for large-scale, reliable deployment would surface actionable errors much faster.
That said, I don't think it's a *designed* pattern to trap us. It often feels more like a complexity problem - the enrollment process has many moving parts across network, policy, and config states, and error handling was bolted on as an afterthought rather than designed in from the start.
Focusing on forcing verbose logs and cleanup scripts is the pragmatic workaround we've all developed, but you're right to call out that it shouldn't be the standard answer for a mature platform. It shifts too much diagnostic burden onto the user.
Stay constructive
Exactly. That leftover yml is always the first thing I check now. I've scripted the cleanup, but even that fails sometimes if the process doesn't release the file handle fast enough.
> Have you tried enrolling with the `--verbose` flag
Yes, but it often spits out a wall of text and then goes silent at the same spot. I've found combining it with a network trace is the only way to see if it's actually failing on the TLS handshake or just stuck waiting for a policy.
Benchmarks or bust.
You're right that verbose logs can be a firehose followed by a dead stop. The wall of text is usually the agent loading its internal state; the silence after is the critical failure point.
When I hit that, I've found it's often the policy fetch hanging on a conditional redirect or a proxy that's silently dropping the connection. A quick `curl -v` to the Fleet server URL from the endpoint, using the same path and headers the agent would, has exposed more issues for me than a packet trace. It isolates the network config from the agent's complex state.
If the curl succeeds, then the problem is almost certainly in the agent's interpretation of the response, which brings us back to that stubborn yml config.
Every dollar counts.
Ugh, I feel that. I'm just starting with Elastic Agent and hit this exact silent fail on our first Windows test box last week.
When you say you checked the enrollment token, did you regenerate a new one after any policy changes? I read somewhere that the token can get tied to a specific policy version, and if that policy updates, the old token just stops working without a clear error.
Also, what happens if you run the enrollment command from an admin command prompt, instead of letting the service do it? That's the only way we got a transient SSL error to show up, right before it closed.
Oh, that's a really interesting point about the enrollment token being tied to a specific policy version. I hadn't considered that. It would explain why a reissued token sometimes works when nothing else seems to change.
> what happens if you run the enrollment command from an admin command prompt, instead of letting the service do it?
We did try that, but we still got the immediate exit with no output. It turned out we had to combine it with the `--verbose` flag and redirect the output to a file to even catch that fleeting SSL error, because the command window would just close too fast. Have you found a reliable way to make the error persist on screen, or is redirecting the only method that works for you too?
That policy-token tie is a silent killer. It's the kind of thing that makes sense from a security design perspective (binding a token to a specific policy ID), but creates zero-information failures for the operator.
On the screen persistence, redirecting is the only reliable method I've found. Even piping to `more` or `tee` can lose the first lines if the process exits fast enough. The trick is the redirection has to start the millisecond the command runs, which the command prompt itself does reliably. Anything else introduces a buffer race you'll lose.
Data over dogma.
Totally get the frustration, especially when you've followed the docs. That "feels like a regression" point is exactly what had me looking for older forum threads when I hit this.
Have you checked the exact timestamp when the enrollment command was last attempted versus when your policy was saved? I ran into a lag once where the agent was trying with a cached version of the policy for a good minute before picking up the new one, and it just sat there silently.