Skip to content
Notifications
Clear all

Help: Client update failed silently and left half our fleet unprotected.

36 Posts
33 Users
0 Reactions
77 Views
(@carlam)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Oof, been there with that "successful" deployment but broken filter state. The blocked outbound connections are the smoking gun.

You've gotten great advice on the CBS.log and CrowdStrike exclusions. I'd add that you should check if Netskope's error is actually logged under a different Event ID in the same CBS channel. We once found the real failure was a few hundred lines *after* the generic exit code.

For a functional test beyond a service restart, we run a quick curl command to their diagnostics endpoint from a non-admin context. If it fails, the user-level agent routing is broken, which a simple service restart won't always fix. You might need a full uninstall/reinstall script for the truly stuck endpoints.


Benchmarking my way to better decisions


   
ReplyQuote
(@alexm82)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Yeah, we've seen this exact thing. Your firewall clue about blocked connections is huge, that's the filter driver dead.

Check the CBS.log like others said, but also look at the Microsoft-Windows-DeviceManagement-Enterprise-Diagnostics-Provider/Admin log. We found a permission error there once that the installer missed.

For CrowdStrike, adding an exclusion for the Netskope installer folder got our update moving. Did you have to adjust your CS policy for this, or did it just start blocking it out of the blue?



   
ReplyQuote
(@coffeelover)
Honorable Member
Joined: 3 months ago
Posts: 397
 

That "false success" is worse than a clean failure. Their management console becomes a liability, telling your team everything's fine while traffic's not filtered.

The race condition you mentioned is real. We solved it by adding a five-minute startup delay to the Netskope service. It loses the race less often. A ridiculous workaround for a vendor's bad driver load logic.


Just my two cents.


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 2 months ago
Posts: 351
 

The "successful" deployment status likely just means the management API call succeeded, not that the installer completed. That's a critical distinction their console fails to make. Your blocked firewall connections are the definitive proof of failure.

On your specific questions:
1. Beyond CBS.log, check the Windows Event Log under `Application and Services Logs > Microsoft > Windows > DriverFrameworks-UserMode`. Look for Event ID 1015 or 2003 from the "User-mode Driver Framework" source around your install time.
2. CrowdStrike's "Containment" or "Prevention" policies are prime suspects. We've seen them flag the driver replacement as a "kernel-level modification" and kill the process. A temporary exclusion for the Netskope service executable (`nskagent.exe` or similar) and the installer .msi is often necessary.
3. The unresolved question is what functional check to implement. A service restart is insufficient if the driver failed to load. You need a script that validates the filter driver is actually attached to the network stack. `netsh wfp show filters` can sometimes show its state, but a real test requires sending traffic through the expected data path and verifying it's tagged.


Trust but verify.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

That's a crucial functional point. The `netsh wfp show filters` command can be a good read, but its output for these types of agents is often obscure.

We've had more success with `sc query type= driver` to confirm the driver is in a "running" state, and then a low-level network test. A simple `Test-NetConnection` or `telnet` to a known-external IP on a filtered port will often hang or fail silently if the filter is stuck, even if the service is reporting as healthy. You need to validate the data plane, not just the control plane.


Less spend, more headroom.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

You're right about validating the data plane being the final word. That `sc query` check is a solid step, but even a "running" driver state can be misleading if it's not actively filtering.

We learned the hard way that some agents will pass a basic outbound test but fail under specific conditions, like UDP traffic or certain ports. Our sanity check now is a small, scheduled script that tries to hit a known-bad IP from a test user's context and logs the result. If the agent is truly healthy, it should block it. It's a bit crude, but it's caught a few "running but blind" filters.


Keep it civil, keep it real.


   
ReplyQuote
Page 3 / 3