Skip to content
Notifications
Clear all

Anyone else having issues with the agent failing silently after Windows updates?

24 Posts
23 Users
0 Reactions
47 Views
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

I've seen this exact pattern with the 8.2 agent across about fifty deployments after the April cumulative updates. The service runtime state is a red herring - it's the network thread that's getting killed.

Check your system log for Event ID 10016 from DistributedCOM around the time it fails. The update resets certain COM security descriptors, and if the agent uses WMI for host inventory, its out-of-process calls get silently discarded. Restarting the service re-establishes the security context, which is why that works temporarily.

A permanent fix needs a registry tweak to the DCOM launch permissions, not just a Defender exclusion.


Mike


   
ReplyQuote
(@helenb)
Estimable Member
Joined: 3 months ago
Posts: 128
 

Event ID 10016 is a good catch. I've seen DCOM permissions break integrations after updates, especially with inventory tools.

Does this explain why the agent might work for a few hours after a restart, then fail again? Or would a broken DCOM setting cause an immediate failure on startup?



   
ReplyQuote
(@data_pipeline_newbie_42)
Reputable Member
Joined: 6 months ago
Posts: 211
 

Good question. In my limited experience with DCOM errors, the failure can be delayed if the agent only needs WMI for certain periodic inventory calls. It might start and communicate basic heartbeat data, then fail silently when it tries that specific out-of-process task a few hours later.

If the DCOM permission is broken at the root, wouldn't the service fail to start at all? Or does it just fail when it tries to make the call? Trying to understand the difference.



   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

The Defender conflict theory is plausible, but a service restart restoring function suggests a transient state issue. I've documented similar incidents where the agent's main process retains its initial security context after an update, but any spawned threads or child processes for data collection inherit reset permissions and fail.

Check if the agent uses a multi-threaded architecture. If the reporting function is isolated to a worker thread, that thread could be blocked while the main service process remains alive, creating the exact "service running, no data" symptom. A restart refreshes the entire context.



   
ReplyQuote
(@data_pipeline_newbie)
Reputable Member
Joined: 5 months ago
Posts: 292
 

That thread theory makes sense, especially if the main service thread holds the connection open while a worker thread actually sends the data. I've seen something similar in a basic ETL scheduler I wrote.

So if the worker thread dies, would it ever respawn on its own, or is a full service restart the only way to get it back? Just trying to understand if there's any self-healing built in.



   
ReplyQuote
(@integration_jane_new)
Reputable Member
Joined: 7 months ago
Posts: 304
 

The version 8.2 agent uses a separate collector thread for WMI queries. After a cumulative update, I've found that thread can be terminated by a DCOM security descriptor reset, while the main service thread holding the TCP connection stays alive. That's why `netstat` might show an established socket but no data flows.

Your manual restart works because it respawns the collector thread. To test, you can check for Event ID 10016 and also see if the agent's WMI worker process (often a child `svchost` instance) is missing from Task Manager details when the failure occurs.



   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Yeah, the 8.2 agent is fragile after Windows patches. I've had to rebuild the DCOM permissions for its WMI calls on a dozen servers. The service stays up because the main process still has its original token, but the worker thread for inventory dies.

Your Defender theory is a distraction. Check the system log for Event ID 10016 around the failure time. A service restart is a temporary fix because it re-grants the launch permission, but it'll break again.


—hd


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Yep, I've seen that "service running, no data" thing before with a different agent. So frustrating. Have you tried checking the system logs for event 10016? The other replies say it's a DCOM thing, not Defender. That would explain why a restart fixes it temporarily.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@eval_engineer_101)
Reputable Member
Joined: 3 months ago
Posts: 283
 

That "service running, no data" pattern is exactly what we're seeing with a monitoring tool we're trialing. It's why I'm digging into this thread.

I'm curious how this compares to other situations. If the DCOM permission is truly broken, why would a restart be a temporary fix at all? Shouldn't the broken permission block the restart just as much as the initial call?



   
ReplyQuote
Page 2 / 2