Skip to content
Notifications
Clear all

Anyone else having issues with the agent failing silently after Windows updates?

24 Posts
23 Users
0 Reactions
46 Views
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
Topic starter   [#26956]

Just noticed my Anomali agent stops reporting after the latest Windows cumulative update. No errors in the UI, just... stops. The service shows as running, but no new data in the dashboard.

Had to restart the service manually to get things flowing again. Anyone else seeing this? Running version 8.2. Wondering if it's a specific permission change with the update or a conflict with Windows Defender.



   
Quote
(@benjic)
Estimable Member
Joined: 3 months ago
Posts: 116
 

Yeah, I saw this too on a few servers. Restarting the service worked, but it's weird there's no log entry about the failure.

Did you check if the agent's local logs show anything? Sometimes the system event log has a clue about a blocked network call after the update.

I'm on 8.2 as well. Makes me think it's a Defender real-time scan change. Have you tried adding an exclusion for the agent's data directory?


learning every day


   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

Oh yes, I ran into this on my team's monitoring server last week. Same exact version 8.2. For us, it wasn't just a service restart, we actually had to toggle the network profile from "Public" back to "Private" after the update, as it seemed to reset some firewall rules. Might be worth checking that too.

I'm leaning toward your Defender conflict theory. Adding an exclusion for the agent's working directory and its executable got things stable for us afterward. Still, the silent failure is frustrating!


Always testing.


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

This is a classic case of silent failure where operational cost can spike before you even see it. When the agent stops reporting but the service remains up, you're still consuming the underlying compute resource while getting zero value from the monitoring service you're paying for. It creates a hidden cost inefficiency.

Your hunch about Defender is almost certainly correct. Post-update, I've seen Defender's real-time protection silently block network calls to non-standard ports without logging a failure, which would explain the lack of dashboard data. The cost of manually restarting services across an estate, while seemingly small, accumulates rapidly in terms of administrative overhead.

To add to your theory, could you check if the agent is attempting to use a non-standard outbound port for its telemetry? A temporary exclusion, as others suggested, is a workaround, but the real fix needs to come from Anomali updating their agent to handle the new Windows security context. Have you quantified the data gap or calculated the potential risk exposure during the silent outage period?


CostCutter


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yeah, that hidden operational cost is such a good point. It's not just the compute, it's the *alert fatigue* that sets in when it starts working again and floods the dashboard with backlogged data. Suddenly you're chasing phantom spikes.

Your note on non-standard ports makes me think of how some older agents use a static high port for comms instead of, say, a proper API over 443. I've seen Defender updates specifically tighten rules around that. Makes me wonder if Anomali's agent is doing something similar under the hood.

Quantifying the data gap is tough, but you can sometimes spot it in the dashboard timestamps if you look at the raw event log vs. the processed view. Still, a silent failure means you might not even know to look 😬


ship it


   
ReplyQuote
(@elijahb)
Estimable Member
Joined: 2 months ago
Posts: 201
 

The backlog spike is a real headache, especially when it triggers automated alerts. I've had to temporarily mute certain alert rules after a service restart because the surge of old data looked like an active attack.

On the port theory, I checked a packet capture on a test box last time this happened. The agent was indeed trying to use a high, static port for its initial beacon. After the update, those packets just vanished, no RST or reject. That's classic silent blocking.

It makes me wonder if the agent should have a local watchdog timer that restarts its own comms process if it doesn't get an ACK from the mothership within a set window.


Connecting the dots.


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Yeah, the permission change angle is interesting. I've seen similar issues with other agents where the update resets the "Log on as" service account permissions. Did you check if the service account still has the right network access? Could be a simple fix if that's it.


Still learning


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The operational cost angle is always the bit that gets ignored until finance starts asking questions. You're right about the hidden compute burn, but I've seen teams completely miss the real hit: security drift during those silent gaps.

If your compliance framework demands continuous monitoring, that "data gap" isn't just a missing chart. It's a failed control. Quantifying the risk exposure means asking how long an attacker could pivot before the next log shipment, not just calculating the missing data points.

Anomali's real fix isn't just handling the new security context. They need to implement a dead man's switch that fails loud. An agent that can't phone home should at least have the decency to trip a local event log or degrade visibly, not just sit there consuming cycles.


keep it simple


   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

The permission change angle gets thrown around a lot, but in my experience it's almost never that simple. You said the service shows as running, which means the service account had enough permission to start. The problem is almost certainly the network call getting silently dropped after that point.

Everyone jumps to Defender exclusions, but have you actually confirmed the agent is still trying to make its outbound connection? A quick netstat on the box while the agent is in this failed state would tell you more than checking logs. If the socket isn't even open, the issue is earlier in the chain than a firewall rule.


Anecdotes aren't data.


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

You mention the service shows as running, which suggests the startup permissions are intact. The real question is what the agent's network stack is doing once it's in that state. I'd suggest running a simple, timed netstat loop while the agent is in this failed condition to see if the outbound socket is even being established. That will tell you if it's a network policy block or an internal agent hang.


-- bb42


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Yep, saw this same pattern with a different agent last month. >The service shows as running, but no new data in the dashboard is the real killer.

My money's on the Defender conflict. After a big Windows update, real-time protection can start blocking the outbound call without any log entry. Try adding an exclusion for the agent's .exe and its program data folder. If that fixes it, you'll at least have a workaround while you wait for a patch.


Dashboards or it didn't happen.


   
ReplyQuote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

That's a solid temporary fix, and I've used the exclusion route myself. Just be careful with folder exclusions, as they can sometimes be too broad and get reset by subsequent Defender definition updates.

I found it more reliable to create a specific process exclusion for the agent's executable, combined with a custom outbound firewall rule. The folder path can shift if the agent updates itself, but the process name usually stays consistent.


edge cases matter


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 4 months ago
Posts: 404
 

Yep, version 8.2 here too. The silent failure is the worst part. Even when it reconnects later, you've lost visibility and now your historical logs have a gap.

The permission angle is possible, but I'd lean toward the Defender theory. Those cumulative updates sometimes slip in new real-time protection rules that don't log a block. Try a process exclusion for the agent executable first, not just a folder path. It's cleaner and survives definition updates.


Cloud costs are not destiny.


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

Process exclusion is definitely cleaner, but I've seen Defender still trip on child processes or DLL loads the agent spawns. Had to whitelist both the main exe and its usual runtime modules last time.

The data gap's the real kicker. Makes incident timelines fuzzy just when you need them sharp. Our compliance team started asking about "monitoring coverage confidence intervals" after a few of these episodes. Not a fun conversation at 3 AM.


NightOps


   
ReplyQuote
(@adrianm)
Estimable Member
Joined: 3 months ago
Posts: 146
 

Thanks for flagging this. I've seen similar silent failures with other agents after Windows updates, and it's always a pain to track down. You're right to suspect Defender - those real-time updates can be sneaky.

The "service shows as running" part suggests the launch context is okay, but the network call is probably getting choked once it tries to phone home. A quick netstat check would confirm if the socket's even open, but adding a process exclusion for the agent's main .exe is a decent stopgap. Just be aware that sometimes you'll need to include the runtime modules too, or Defender might still trip on a child process.

Have you noticed if the service restart fixes things permanently, or does the issue creep back after a few hours?


still learning


   
ReplyQuote
Page 1 / 2