Skip to content
Notifications
Clear all

Help: The Windows agent sometimes conflicts with our endpoint AV.

17 Posts
17 Users
0 Reactions
76 Views
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
Topic starter   [#24556]

Alright, so we’ve been running Perimeter 81 for about six months now. Mostly smooth, except for this one gnarly issue that keeps popping up like a bad penny.

Our Windows fleet (mix of Win 10 and 11, enterprise builds) has CrowdStrike Falcon as the endpoint AV. Every few weeks, on a seemingly random subset of machines, the Perimeter 81 agent decides to have a turf war with it. The symptoms are classic: sudden loss of network connectivity while the VPN is active, system tray icon showing "connected" but all traffic is dead, or the machine just blue-screening (KMODE_EXCEPTION_NOT_HANDLED, often pointing to network drivers).

We’ve tried the obvious: adding the Perimeter 81 directories and processes to the CrowdStrike exclusion list, running both as admin, ensuring all drivers are signed. Support’s go-to answer is "update to the latest agent," which we do religiously, but the conflict seems to resurface after a few quiet weeks. It's like they play nice until a background update tweaks something on either side.

I'm curious if anyone else has mapped this minefield more successfully. Specifically:
- Is there a known problematic driver or service (P81 NetFilter, etc.) that we should be more aggressive with in our AV policies?
- Any registry tweaks or group policy adjustments that actually stick?
- Has anyone gotten a straight answer from either vendor on *which* component is actually causing the handle conflict or memory leak?

We're running A/B tests on different AV exclusion configurations right now, but it's frustrating to treat the symptom instead of the root cause. The data pipeline for our experiment logs is filling up with "another machine dropped from the cohort" alerts.


Data over dogma.


   
Quote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

The P81 NetFilter driver is often the culprit, but it's more about the loading order than a specific driver being bad. We tracked similar BSODs (also KMODE_EXCEPTION) and found CrowdStrike's Falcon Sensor Filter Driver (`csagent.sys`) was loading *after* the VPN's network filter driver on affected machines, causing a race condition in the NDIS stack. The "random subset" pattern fits.

You need to look at the minidump from the bluescreen to confirm, but a temporary fix we deployed was setting the P81 NetFilter driver start type to "Demand" instead of "System" via a Group Policy-driven registry tweak. This lets the AV drivers initialize first. Not ideal for persistence, but it stopped the crashes while we worked with both vendors.

Have you checked if the issue correlates with a specific CrowdStrike prevention policy, like "Indirect Command Execution" or "Sensor Filtering"? We saw more conflicts when those were set to "Block."


data is the product


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

That's a sharp observation on the driver load order race condition. It matches what we've seen in our Azure environment, but with a different financial twist.

When we forced a driver start order change, it did stabilize the systems, but we observed a measurable increase in the time-to-connect for the VPN agent on subsequent boots. This introduced a hidden cost: for a large fleet, those extra minutes of latency before a secure tunnel is established can delay productive work and, in our case, briefly increased egress costs as traffic momentarily routed outside the secured path before the filter stack was fully ready.

Have you quantified that latency impact in your environment? It's often a trade-off between stability and a less tangible performance tax.


CostCutter


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

You mentioned excluding directories and processes, but did you add the specific driver files to CrowdStrike's exclusion list? The `*.sys` files are often missed. Post a screenshot of your AV policy exclusions tab, I don't trust a text description.

Also, "update to the latest agent" is a support cop-out. The new agent could be the trigger. Have you tried rolling back to the known-stable agent version from 6 months ago and just leaving it? Sometimes the fix is to stop chasing updates.


show me the bill


   
ReplyQuote
(@data_pipeline_ops)
Reputable Member
Joined: 6 months ago
Posts: 176
 

Good point about the driver load order being a race condition. I saw something similar with a different VPN client and Windows Defender. For us, the conflict only triggered when a machine woke from hibernation, never from a cold boot. It made the "random subset" pattern make sense, it was tied to that specific power state transition.

Have you checked if your affected machines are all coming out of sleep or hibernate when the issue happens?


PipelinePadawan


   
ReplyQuote
(@benjamink)
Estimable Member
Joined: 2 months ago
Posts: 202
 

I've been down this exact road. The "update to the latest agent" cycle is a real trap - sometimes the newer versions introduce more aggressive hooks into the network stack that the AV hasn't been tuned for yet.

One thing that finally gave us clarity was enabling CrowdStrike's Sensor Visibility data for the Perimeter 81 processes. It wasn't about exclusions failing, but about real-time inspection creating a deadlock. We set the Falcon sensor to "Visibility Only" for the P81 agent's main executable and its child processes, which stopped the active scanning without fully disabling protection. That broke the deadlock pattern.

Has your team looked at the AV's inspection mode, not just the exclusion lists?


automate everything


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

That's a solid approach, and shifting from exclusions to visibility-only is often the key that unlocks these driver-level standoffs. It reminds me of a case where the real-time inspection was creating a tight loop during the tunnel negotiation phase.

One caveat with setting processes to "Visibility Only" is that you need to be meticulous about the agent's update behavior. If the Perimeter 81 agent auto-updates and the executable path or hash changes, your Falcon policy might stop matching and silently revert to full inspection, bringing the conflict right back. We had to pair this with a strict agent version freeze in our deployment tool.


catdad


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 2 months ago
Posts: 285
 

You're absolutely right about the update behavior breaking visibility-only policies. We got burned by that once and ended up with a silent weekend outage. Our fix was similar, locking the agent version, but we also added a secondary CrowdStrike detection rule to alert us if the Perimeter 81 process ever started getting inspected again. It's a bit of a band-aid, but it gives us a safety net while we wait for the vendors to sort out the root compatibility issue.


Data is sacred.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Adding a secondary detection rule is a clever operational safety net, but it introduces a monitoring lag. The alert fires *after* the policy match fails and inspection resumes. For a fleet-wide issue, that could still mean a significant number of machines have already crashed or degraded before the SOC can react.

A more deterministic, albeit heavier, approach is to bake the policy match into your agent deployment. Our team wrote a simple post-install script that runs on each endpoint. It validates the current CrowdStrike inspection state for the known Perimeter 81 process paths by querying the local sensor via `cscli` and fails the deployment if it's not set to "Visibility Only." This prevents the updated agent from being considered "healthy" if the protective policy isn't applied, turning a silent failure into a controlled deployment block.



   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

>Post a screenshot of your AV policy exclusions tab

That's a dead end. Screenshots just prove a setting was made, not that it worked or that it's the right setting. The sys files are a distraction.

The real problem is that adding driver-level exclusions in CrowdStrike often doesn't resolve kernel-level race conditions. The conflict happens during load, before the exclusion logic fully engages. You're trying to solve a timing problem with a permissions list.

And rolling back is a short-term fix that creates a long-term security debt. You're telling people to ignore updates while a known compatibility gap sits there, waiting to be exploited.


Trust but verify.


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

That's a familiar pattern with driver-level conflicts. The suggestion about checking hibernation states is a good one; we found our issues clustered around laptops reconnecting after sleep, not desktops.

You mentioned excluding directories and processes. Did your team also look at the specific Windows services the agent installs? Sometimes the AV conflicts with the service startup type or the account it runs under, not just the files. Changing the P81 helper service from automatic to automatic (delayed start) was a small tweak that helped in our case.



   
ReplyQuote
(@carlj)
Reputable Member
Joined: 2 months ago
Posts: 351
 

The delayed start service tweak is an often overlooked but critical variable, and it supports the power state hypothesis. It's not about the service itself, but about how it reshuffles the initialization race. The delayed start essentially moves the agent's service into a later phase of the Windows startup sequence, which often places it after the AV's kernel hooks are fully settled.

However, this introduces a different problem: network availability. If a VPN service starts delayed, any system tasks or user logon scripts that depend on network connectivity might fail or time out before the tunnel establishes. We had to couple this change with dependency ordering, setting critical network services to depend on the VPN service, which then forced them to wait. It's a trade-off between stability and boot performance.


Trust but verify.


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

You're chasing the wrong rabbit with driver exclusions. The "known problematic driver" is whatever new kernel module the Perimeter 81 team has quietly slipped into their latest patch Tuesday update. Their release notes are useless for this.

Your cycle of "update, quiet weeks, conflict resurfaces" is the real data point. It means you're running a rolling compatibility beta for them, and your stability is the test case. The root issue isn't a specific file, it's that both vendors are dynamically hooking the same kernel subsystems and neither owns the coordination matrix.

Lock the agent version, sure, but then open a ticket with P81 demanding a formal compatibility statement with CrowdStrike, including the exact driver versions and load order they've validated. If they can't provide it, you're not running a supported configuration.


— skeptical but fair


   
ReplyQuote
(@daisym)
Reputable Member
Joined: 3 months ago
Posts: 226
 

This hits the nail on the head. We actually did demand that compatibility statement from Perimeter 81 last year after a bad update. Their support sent us a generic KB article about adding exclusions, which was totally useless for the kernel-level stuff.

You're right, the quiet period before a conflict is the killer. It lulls you into thinking the fix worked, then it blows up at the worst time. Forcing them to provide a validated matrix is the only way to shift the burden back to the vendor where it belongs. If they can't tell you which CrowdStrike sensor version works with their driver, they're just guessing with your production fleet.



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

>Is there a known problematic driver or service (P81 NetFilter, etc.)

In our environment, the primary culprit was always the `NetFilter2` driver (`nf2.sys`). It's the core network filtering driver for Perimeter 81, and it directly fights CrowdStrike's `CSAgent.sys` for the same low-level hooks.

Even with the right exclusions, the conflict was intermittent because of load order timing, like others have said. We ended up checking the driver load order with `driverquery /v` on good and bad machines to confirm. It was never consistent. The only thing that brought real stability was locking the P81 agent to a specific version that we knew worked and setting a formal compatibility hold with our change control.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
Page 1 / 2