Skip to content
Notifications
Clear all

Help: The 'always-on' feature on Windows causes a login loop on our domain-joined machines.

21 Posts
21 Users
0 Reactions
5 Views
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
Topic starter   [#29135]

Hey folks, ran into a real head-scratcher this week and thought I'd bring it to the community. We've been rolling out NordLayer to our data engineering team for secure access to our cloud data lakes (Snowflake, S3 buckets, you know the drill). On personal machines, it's been smooth sailing.

But on our corporate, domain-joined Windows 10/11 machines, we're hitting a nasty loop when the "always-on" VPN feature is enabled. The setup: machine boots, NordLayer auto-connects, but then it seems to interfere with the domain login process itself. The user gets to the login screen, enters credentials, the screen flashes, and it drops them right back to the login screen. Only workaround so far is safe mode to disable the client or the feature.

Feels like a routing conflict or maybe a credential passthrough issue? The domain controllers are on-prem, and I'm wondering if the VPN tunnel is blocking the necessary authentication ports (like 445, 88, 389) before the user session is even established. Has anyone else deployed NordLayer in a similar AD environment?

Our config is pretty standard:
```json
We're using the "allow local network access" option,
with split tunneling OFF, hoping to route all traffic through the tunnel.
Protocol is set to NordLynx (WireGuard).
```

Would love to hear if you've found a magic bullet. Maybe a specific order of GPOs, a registry tweak, or a known incompatibility. We need this "always-on" for compliance, but the login loop is a showstopper for our automated pipelines that rely on domain service accounts.

ship it


ship it


   
Quote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Yeah, the split tunneling off is likely the culprit. You're routing *all* traffic, including the pre-login authentication chatter with your domain controllers, through the VPN tunnel. If that tunnel isn't fully established or the route back to your on-prem network is blocked, the machine can't complete the login.

Had a similar fight with another client VPN last year. The fix was creating a very specific routing rule in the VPN config to exclude our internal AD subnets (the DCs, DNS, the whole shebang) from the tunnel. But that depends if NordLayer's admin console gives you that granularity. Might need to involve their support to see if they can push a policy excluding your internal IP ranges.



   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

You're definitely on the right track thinking about authentication ports being blocked. >The domain controllers are on-prem< - that's the core issue. When split tunneling is off, the VPN tunnel becomes the default route before any user logs in, so the machine can't reach the domain controllers on ports 88 (Kerberos) or 389 (LDAP).

We hit the same wall with a different always-on VPN. The fix involved adding persistent static routes *before* the VPN client loads at boot. You can push this via a startup script or GPO, something like:
```cmd
route -p add 10.0.0.0 mask 255.255.0.0 10.0.0.1
```
...where you're routing your internal AD subnet through the local gateway, not the VPN tunnel.

Check if NordLayer's admin console has an "exclude local subnets" or "split tunneling" policy you can push. If not, those static routes are your next best bet before login.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 2 months ago
Posts: 308
 

That description of the login screen flashing and dropping back is textbook. It's exactly what happens when the machine can't talk to a domain controller at that precise moment.

You've hit the core issue by saying "the domain controllers are on-prem." With split tunneling off, the VPN tunnel becomes the active network interface before user login. All traffic, including the authentication request to your on-prem DCs, tries to go out the tunnel, and that just can't work.

I'd start by checking if there's an "allow local network access" or split tunneling setting in the NordLayer admin portal that you can push via policy. You need to exclude your internal AD subnets so that traffic stays on the local network. If their admin console doesn't offer that granularity, you'll likely need to get their support involved, as manually adding static routes is a workaround, not a sustainable policy.


Reviews build trust.


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 2 months ago
Posts: 349
 

Ah, that "allow local network access" option with split tunneling OFF is the key detail! That combination is causing your headache.

I've seen similar loops with other always-on VPNs. The setting sounds like it should work, but sometimes the VPN client's route metric still takes precedence, choking off the DC connection. You might need to test it with split tunneling *ON* and only send the specific cloud ranges through the tunnel. That often gives the local gateway priority for internal traffic.

Have you checked the exact routing table on a test machine right after boot, before login? That could confirm if the route to your AD subnets is still pointing to the local network.


null


   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

You're staring straight at the problem, but you're accepting the vendor's checkbox as a solution. The "allow local network access" option with split tunneling OFF is a complete gamble and often a lie. It suggests the client will *try* to route local traffic locally, but route metrics and timing during boot are finicky. The VPN tunnel's interface often still gets a lower metric, winning the race and hijacking the DC traffic.

Don't trust the checkbox. Run `route print` from the login screen (you can launch cmd via accessibility tools) and see where 10.0.0.0/8 (or whatever your internal range is) is actually pointing. I'll bet a coffee it's using the VPN gateway.

The real fix is forcing the issue: use GPO to push persistent static routes for your AD subnets with a lower metric than the VPN, or ditch this config entirely. Go split tunneling ON and only tunnel the specific cloud IP ranges you need. It's more work, but it works. This "best of both worlds" config they're selling usually means "neither world works reliably."


pay for what you use, not what you reserve


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

Exactly, the split tunneling being off is the key. When you said >similar fight with another client VPN< it reminded me of our own saga with Zscaler's always-on client a while back.

We got caught in the same loop, and the admin portal had a checkbox for "exclude corporate subnets," which sounds perfect. But we learned the hard way that you have to define those subnets *incredibly* precisely. We initially just excluded our main office subnet, forgetting our secondary DC in a different /24 at a colo. The login would work at HQ but fail for remote-office machines trying to auth with that backup DC. Had to meticulously list every single internal IP range that could host a domain controller or DNS server.


hannah


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

You've nailed a critical detail that's easy to miss. That "exclude corporate subnets" feature is only as good as the list you put in it. It's not just domain controller IPs either, you have to consider all the services that feed into the authentication handshake, like your DNS servers and maybe even your WSUS servers if they're involved in policy application. Missing one /24 can leave a whole group of users stuck.

This is where a good network diagram, or at least a consolidated list of all internal service subnets, becomes essential before you even start configuring the VPN policy.


Keep it constructive.


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

You're absolutely right about the authentication traffic trying and failing to go out the VPN tunnel. That's the mechanism of the failure. The operational challenge, which I think you've hinted at, is the sheer number of potential destinations that need to be excluded. It's rarely just the DCs.

For a sustainable policy, you need a meticulous inventory that includes:
* All domain controller IPs/subnets
* DNS servers (critical for DC location)
* Possibly even time servers (NTP) if they're internal
Missing any one of these can cause subtle, location-dependent failures that are a nightmare to debug. The vendor's admin portal might only let you exclude, say, 10 subnets, which could be insufficient for a large, distributed organization. That's when the "static routes via GPO" workaround becomes the only scalable policy.


CostCutter


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

Oh, you've hit on the exact frustration. That "exclude subnets" list becomes a monster. We had to maintain a spreadsheet just for this, and it was a huge pain when a new server subnet was spun up and forgotten.

One caveat to your great list: don't forget any internal CA servers if you're using certificate-based auth for anything. Left that off our initial list once and it broke a very specific set of laptop logins.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 2 months ago
Posts: 342
 

That "allow local network access" with split tunneling off is the exact combo that's probably giving you grief. I've chased a similar ghost with another client VPN where that setting created a false sense of security.

You mentioned a pretty standard config, but have you checked the routing table immediately after boot, from the login screen itself? You can launch a command prompt by using the accessibility icon (narrator) and checking `route print`. I'm willing to bet your internal AD subnet is still being routed through the VPN gateway's interface because its metric is lower. The checkbox often doesn't adjust the metric aggressively enough.

If that's the case, you might need to bypass the checkbox entirely and use a startup script to add a persistent, high-priority route for your domain controllers before the VPN service fully initializes. Something like `route -p add mask metric 1`. It's a bit manual, but it wrestles back control from the VPN client's routing logic.


null


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

You're focusing on the right symptoms, but you've got a fundamental misunderstanding in your config. The combination of "allow local network access" with split tunneling OFF is inherently contradictory - it's asking the VPN client to both capture all traffic *and* let local traffic bypass it, which creates a routing race condition.

The login screen flashing back is the telltale sign that authentication traffic to your on-prem DCs is taking the VPN tunnel and failing. Before you even check ports, check the route table immediately after boot with `route print` launched via the accessibility menu. I'd wager your internal subnets show the VPN gateway as the next hop.

You need to either:
- Enable split tunneling properly and define only your cloud data lake CIDR ranges
- Or abandon that checkbox and implement forced local routes via GPO startup script

Your current configuration is creating an unrecoverable network state at boot.


infrastructure is code


   
ReplyQuote
(@annak8)
Estimable Member
Joined: 2 months ago
Posts: 202
 

That inventory you've laid out is spot on, but man, is it a painful process to get right. We went through the exact same exercise, and even with a list that felt exhaustive, we still missed a critical piece: the DFS namespace servers hosting our SYSVOL. The login process tried to reach out to them for policy files, hit the tunnel, and timed out. It wasn't an immediate loop - it just added a weird 90-second delay to login that drove everyone nuts until we traced it.

It really does become a spreadsheet nightmare, like user1526 said, and you're forever worried about a new server subnet getting stood up without anyone updating the VPN exclusion list. Makes you appreciate the static route GPO method, even if it's clunkier to set up initially.



   
ReplyQuote
(@connork)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Yeah, that sounds exactly like what happened to us with a different client. The "allow local network" option didn't save us either.

We found the login screen can't reach the domain controller because all the traffic tries to go out the VPN tunnel. Like others said, the checkbox seems like a lie sometimes.

How detailed is your "local network" definition in NordLayer? Did you list every single subnet your DCs and internal DNS live on? Missing even one little subnet can cause this loop.



   
ReplyQuote
(@ava23)
Honorable Member
Joined: 2 months ago
Posts: 435
 

Ugh, the classic "checkbox that doesn't really do what it says" in vendor land. You've put your finger on the exact problem everyone's dancing around: the VPN is hijacking the auth traffic.

You said it feels like a routing conflict? That's exactly what it is. The "allow local network access" with split tunneling OFF is a vendor fantasy. It tells the client, "capture all traffic, but also, don't capture this traffic." The client's routing logic panics, your DC subnet gets a route via the VPN interface with a better metric, and goodbye login.

Checking ports is a red herring at this point. The packets aren't even getting a chance to try the right ports because they're sent down the tunnel to a gateway that has no idea how to get back to your on-prem DCs.

Before you go down the spreadsheet-of-doom exclusion list rabbit hole, check the actual routing table from the login screen like others said. I bet it's a mess.


Trust but verify.


   
ReplyQuote
Page 1 / 2