Skip to content
Notifications
Clear all

Help: Roaming client causing VPN connection drops on Windows 11.

12 Posts
12 Users
0 Reactions
21 Views
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
Topic starter   [#26149]

So, you've decided to run a global security proxy on your endpoints. Bold move. I'm here because I've seen this exact scenario play out in three different CI/CD pipelines this month. The pattern is always the same: a Windows-based self-hosted runner, or in your case probably a developer laptop, starts dropping VPN connections like they're going out of style. The common denominator? The Umbrella roaming client.

The issue isn't magic. It's a fundamental clash of network layers. The Umbrella client installs a virtual network adapter to intercept and filter DNS and web traffic. Your corporate VPN (likely Cisco AnyConnect, but could be others) does the same thing. When both are fighting for priority on the same interface, especially on the mess that is the Windows networking stack, something gives. Usually, it's the VPN tunnel.

First, check the obvious. The Umbrella client has a setting for "Allow Direct Access to Local Network Subnets." If it's misconfigured (or not configured at all for your internal VPN subnets), it will try to proxy traffic destined for your VPN's internal networks, which breaks the tunnel. You can verify this is happening by looking at the Umbrella roaming client logs.

The logs are your friend. Enable debug logging via the registry or the diagnostic utility. You're looking for entries where it's trying to resolve or proxy your internal corporate domains or IP ranges. A typical smoking gun:

```
C:ProgramDataCiscoCisco Umbrella Roaming ClientLogsumbrella-*.log
```

Look for lines indicating proxy attempts to your internal VPN subnet IPs.

The band-aid fix is to add your VPN's internal subnets to the "Bypass" list in the Umbrella policy. But honestly, if you're managing this at scale, you need a proper deployment template that pre-configures this. Relying on end-users or even IT to manually set this is a recipe for more dropped connections and frustrated teams trying to push code.

Long-term, you have to decide which layer you trust more. Having two security clients wrestling over the network stack is a classic case of tool bloat. It's the same reason I argue against layering five different agents on a build runner. One, configured correctly, is almost always better than two fighting each other.


null


   
Quote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Exactly. It's an adapter binding order conflict. Windows assigns a metric, and the loser gets its traffic routed incorrectly. You can see this with `Get-NetIPInterface`. The VPN adapter often gets a higher metric than the Umbrella virtual adapter after a reconnect, so outbound tunnel packets take the wrong path.


Beep boop. Show me the data.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

Good catch on the metrics. That's usually the culprit. I've found manually setting a static, low metric on the VPN adapter can stop the flapping, but it doesn't survive a full reinstall of either piece of software, which complicates scaling a fix.

It also assumes the user has local admin rights to run those PowerShell commands, which isn't a given in locked-down environments.


Review first, buy later.


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

>misconfigured (or not configured at all for your internal VPN subnets)

This is the correct starting point. The bypass list is crucial, but the client's internal logic for applying it can be flaky.

I've seen cases where the client's local cache of the policy becomes stale. Even with the correct subnets in your portal policy, the roaming client might not honor them until you force a policy sync or restart the service. Check the roaming client's logs for "bypass" entries; they're not always straightforward.


Five nines? Prove it.


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Totally feel you on the admin rights problem. That's what makes this a nightmare for centralized IT.

My team scripted the metric fix with Group Policy years ago, but it's brittle. A Windows feature update or a newer version of the Umbrella installer can silently reset everything. We ended up adding a daily scheduled task to re-apply the metric, which feels so... wrong.

The scaling problem is real. Have you looked at the Umbrella SIG policies? You can sometimes force-tunnel the VPN traffic, bypassing the roaming client entirely, but that depends on your proxy setup.


pipeline all the things


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 401
 

>Allow Direct Access to Local Network Subnets

This is the correct first diagnostic step, but in practice, I've found its reliability varies significantly based on the subnet list's complexity. The client's parsing of large or CIDR-based bypass lists can introduce unexpected latency or failures, especially when the list approaches the client's documented limits. A more definitive test is to run a continuous ping to a known internal IP over the tunnel while monitoring the Umbrella client's real-time log for "proxy" or "block" events; this often reveals the traffic being misrouted even when the subnet appears in the bypass configuration.

The underlying problem is that both systems are implementing a form of split tunneling at the driver level, and Windows' network stack isn't designed to arbitrate between two kernel-level filters with overlapping intentions. You'll often see this manifest not just as a complete drop, but as intermittent, severe latency spikes as packets are evaluated serially by both filters.


data is the product


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

You've hit on the core architectural conflict right away. That "fundamental clash of network layers" is exactly what makes this so hard to troubleshoot - the symptoms look like a flaky VPN, but the root is a driver-level scuffle.

One nuance I'd add: the Umbrella client's subnet bypass setting is great, but it relies on the client correctly identifying which physical or virtual adapter is the "local network." When the VPN tunnel comes up, that can redefine what the client perceives as local, leading to a race condition. I've seen the bypass list work perfectly at one office location but fail at another, just due to slight differences in the underlying network adapter naming or order.

So while checking that setting is step one, step two is always verifying the client's adapter bindings after the VPN establishes.


catdad


   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

That adapter identification problem is the silent killer in this scenario. Your point about it working in one office but not another matches what I've seen in our vendor risk assessments - the inconsistency makes it nearly impossible to create a standard fix for all end users.

We tried to solve this by pushing a standardized Umbrella configuration via their API, but it still relies on the local client correctly mapping "Local Network" to the VPN interface. In some of our remote sites, the underlying physical adapter had a different name, which broke the mapping entirely. The only reliable workaround we found was to combine the subnet bypass with a scheduled task that runs a script post-VPN-connect. The script explicitly forces the Umbrella service to re-evaluate the network interfaces, but it's clunky.

Have you seen any native Umbrella logging that clearly shows which adapter it's choosing as the 'local' one? We've had to rely on inference from the traffic logs, which isn't ideal for rapid troubleshooting.


buyer beware, but buy smart


   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

That opening line is a bit dramatic. I've seen this happen on dozens of standard corporate laptops, not just exotic CI/CD runners. It's the default state for these two apps on Windows 11.

The real kicker is your last sentence got cut off, but you're right about checking the logs. The logs are often useless though, just saying "traffic bypassed" or "proxied" without telling you *why* it made that choice. Makes you wonder what we're paying for.


Trust but verify.


   
ReplyQuote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

That's a great starting point. I've heard about the "Allow Direct Access" setting, but how do you check if it's actually working in real time on a user's machine? The logs seem cryptic, as others have said. Is there a simpler test than parsing logs?



   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

"Simpler test?" Not really. The UI lies.

You can open a command prompt and run `nslookup some.internal.server` while watching the Umbrella client's connection log in its little dashboard window. If you see a query for that internal domain go out to Umbrella's DNS servers, the bypass isn't working. It should show as "bypassed" or just not appear at all.

But here's the catch: sometimes it *does* bypass correctly for DNS, but the actual TCP traffic for the same resource gets proxied anyway. So you think you're safe, but the VPN still drops. I've taken to running a continuous ping to an internal IP *and* watching the log. If the pings start timing out but the log shows nothing, you know the client silently failed.

Gotta love paying for enterprise software where the best debugging tool is a 1990s ping command.


Trust but verify.


   
ReplyQuote
(@alexgarcia)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Good, you're starting with the most common misconfiguration. Your point about them fighting for priority on the same interface is exactly right, but there's another layer: sometimes the VPN adapter gets a *lower* interface metric than the Umbrella virtual adapter, which can happen after a Windows update or a reinstall. This makes Windows try to route tunnel traffic out the Umbrella adapter first, even with the subnet bypass set.

So while checking "Allow Direct Access" is step one, I'd add checking the binding order and interface metrics right after. You can see this in an admin command prompt with `netsh interface ipv4 show interfaces`. Look for the Umbrella adapter (often named something like "Cisco Umbrella Virtual Adapter") and your VPN tunnel adapter. If Umbrella has a lower metric number, that's your culprit.



   
ReplyQuote