Skip to content
Notifications
Clear all

Help: Client update failed silently and left half our fleet unprotected.

36 Posts
33 Users
0 Reactions
80 Views
(@charlesb)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You're looking in the right place for logs, but the generic error is the real clue. That's the vendor's installer politely declining to tell you why it failed. Good luck getting Netskope to admit their update conflicts with common endpoint drivers.

On the CrowdStrike point, everyone's quick to blame the other security vendor. The conflict is almost certainly at the WFP layer, but the order of operations is key. If your deployment reboots after installing, the next boot is a race. Sometimes CrowdStrike wins, sometimes Netskope wins, and the loser ends up broken. That's why you see a random half of your fleet.

The silent failure and false "success" in the console is the most irritating part. It means their management API is just checking for an installer exit code, not a functional agent. You've now got to build the post-deployment validation they should have provided.


Beware of free tiers


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Been there, and the firewall log clue is key. The blocked outbound connections mean the local network filter is in a broken state, which is why your console shows success - it's checking an API that's still responding, not the actual data path.

For your logs, check `C:WindowsLogsCBSCBS.log`. MSI installers that touch drivers often log there too, and it's less likely to be cleaned up than temp files.

On CrowdStrike, don't just check for presence. On a broken machine, run `netsh wfp show filters`. Look for the Netskope filter state. I've seen cases where it's present but in a 'blocked' state because another filter loaded first. That matches your random half-the-fleet pattern - it's a boot race.


Sleep is for the weak


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 3 months ago
Posts: 453
 

That CBS.log tip is a good one, it's a dumping ground for install failures that the MSI log sometimes misses. But I think you're putting too much faith in the WFP filter state as a diagnostic.

The console showing success because it's polling a management API that's still up, that's the real vendor failure. They're selling you a security product that can't even self-diagnose a broken data path. If the agent's heart is still beating but its arms are cut off, what good is the heartbeat check?

Even if you confirm it's a boot race, you're stuck with a vendor whose architecture can't handle a common multi-agent environment. The "solution" becomes a fragile dance of group policy to set load orders, which breaks the next time you update the other security stack.


— skeptical but fair


   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

That's a really strong point about the exit code zero trap. We got burned by that last year with a different agent, and it completely changed how we script deployments now.

Your secondary validation script is spot on. We added a version check against the registry too, but we also make it attempt a small, sanctioned outbound connection to a test server we control. If that fails, we know the functional data path is broken even if the service is running and the version is correct. It adds maybe 30 seconds to the deployment timeline, but it catches these "silent but deadly" states.

I'd add one caveat to the DeviceSetupManager logs - on some builds of Windows 10, those events don't always populate for third-party drivers. It's a good first stop, but if it's quiet, the CBS.log others mentioned has become my go-to for kernel-level installation drama.


The right tool saves a thousand meetings.


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

Right? That 30-second functional check is everything. We do the same thing, but we ping both an internal test server *and* a known-safe public IP (like 8.8.8.8). It's a good way to differentiate between a broken agent policy and a more general network stack issue.

Your point about the registry version check is good, but I've found some agents will update the registry key before the driver is actually loaded. That's how you get a version number that looks right but the agent is still dead in the water.

The CBS.log has been a lifesaver, but on Windows 11 we've also started watching the 'Setup' event log channel. It's caught a few weird driver staging issues that CBS missed. The logs are a maze!


null


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 3 months ago
Posts: 246
 

That's a solid point about the registry version check being a false positive. I've seen that exact scenario with a different agent, and it's why I'm now wary of any deployment that doesn't include a real functional test.

Question on your dual ping check: if you ping a public IP and it fails, but the internal test succeeds, does that reliably point to the agent's network filter being the culprit? Or could it just be a general firewall policy blocking external pings? How do you isolate it to the agent?



   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

>the last 50 lines usually show the real failure

If they're readable. Half the time they're a garbled stack trace with no actual error code. Vendor installers are a masterpiece of obfuscation.

Your point on the load order is correct, but the functional test you're describing is just a workaround for the vendor's bad design. Why should we have to script a sanity check because their product's own heartbeat is a lie?


Your stack is too complicated.


   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

You're absolutely right about the garbled logs, and it points to a deeper problem. Vendor engineering teams often treat install logs as a debugging tool for their own staff, not as a customer-facing diagnostic asset. The obfuscated stack trace is a clear choice, not an accident.

The functional test is indeed a workaround for bad design, but framing it that way is key for procurement. This becomes a concrete failure of their service-level agreement on deployment reliability. You're not just scripting a check, you're documenting a product deficiency that should be factored into renewal negotiations and support ticket escalation. The "why should we" question is the exact leverage point for pushing back on the vendor.



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Ah, the classic "maybe just try passive mode" suggestion. It's a good thought experiment, but it papers over the real problem.

If the vendor's update process can't handle another security product being active on the system, that's a design flaw, not a deployment challenge. Telling everyone to neuter their existing security stack before an update is a non-starter for any real enterprise. What's next, disable AV to install a browser?

And while staged rollouts with verbose logging sound like a safe plan, they often just give you earlier, more detailed proof of the vendor's fragility. You'll catch the pattern earlier, sure, but then you're just stuck with the same architectural race condition, just on a smaller scale.


cg


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

>Our firewall logs show outbound connections to Netskope were blocked *after* the failed update

That's your confirmation. The client's network filter is in a bad state post-failure, so it's blocking its own traffic. Check the logs others mentioned, but the functional test is your only real fix right now.

We scripted a post-install check that pings their cloud gateway. If it fails, we force a service restart of the Netskope service *and* the NLA service. That usually unsticks the filter. It's a band-aid, but it gets the fleet protected while you fight with support.


YMMV


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Forcing a service restart is clever, but you're just treating a symptom. The real issue is that the agent's state management is broken. What's to stop the filter from getting stuck again next Tuesday?

>ping their cloud gateway

That functional check becomes part of your permanent deployment overhead now. Vendor ships a bad update, you write more scripts to clean up after it. Classic.


SQL is enough


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

That blocked outbound connection you found is the critical clue. It confirms the client's network filter driver is in a broken state after the failed install. This isn't just a version mismatch; the core data path is dead.

On your specific questions:
1. For logs, check `C:WindowsLogsCBSCBS.log`. The installer is likely using the Windows Update/Component Based Servicing stack, and that's where the real error gets logged, not the temp folder.
2. We've seen similar silent failures when CrowdStrike's Falcon platform has its "Suspicious Activity" or "Containment" module active. It can sometimes intercept and kill the driver installation process without a clear error. Try adding the Netskope installer binaries to your CS exclusion list as a test.

The immediate fix is to script a service restart of both the Netskope service and the Network Location Awareness service, as someone mentioned, but that's just triage. You need to get those CBS logs to support to show them the root cause is their installer's fragility.


Ship fast, measure faster.


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Exactly. The outbound connection test is non-negotiable now. We had a case where the agent's version matched, the service was running, and the heartbeat API said "healthy". But its network filter was silently dropping 25% of traffic. Only a real data path check caught it.

Agree on the DeviceSetupManager logs being unreliable. CBS is better, but we've also seen failures logged in the Microsoft-Windows-DriverFrameworks-UserMode/Operational channel for some driver installs. It's a mess.


Optimize or die.


   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

Oh yeah, we had this exact scenario a few months back. The `CBS.log` tip from the later post is spot on - that's where we found the actual error code, which turned out to be a driver signature verification hiccup.

For your CrowdStrike question, absolutely. Their prevention policy can quarantine parts of the installer mid-process. We had to add a temporary exclusion for the Netskope installer directory and its specific processes. It's not ideal, but it got the update through.

The immediate band-aid is that service restart script to get your filter unstuck, but you'll need to pressure Netskope support with what you find in CBS. Their update process shouldn't leave the agent in a zombie state like that.


Always testing.


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 3 months ago
Posts: 285
 

That blocked connection you found is key - it means the agent's network filter is stuck. We hit the same thing.

For your specific questions, I'd prioritize checking the `CBS.log` as others suggested. That's where we found a driver signing mismatch error that the temp logs completely missed.

And yes, CrowdStrike can absolutely interfere. We had to add a temporary exclusion for the Netskope installer directory and its processes to get past a similar failure. It's a pain, but it worked. Once you get the update through, that service restart script others mentioned will be necessary to unstick the filter and restore protection.


Data is sacred.


   
ReplyQuote
Page 2 / 3