Right, you've hit the classic vendor checkbox trap. Your standard config isn't standard, it's broken. The "allow local network access" with split tunneling OFF is a logical impossibility the client can't resolve, so it defaults to hijacking all traffic, including your AD auth.
Forget checking ports. The packets aren't even trying the right path. Do this right now: boot a test machine to the login screen, use the accessibility shortcut to launch a command prompt, and run `route print`. I guarantee you'll see your internal corporate subnets listed with the VPN gateway's interface as the next hop and a suspiciously attractive metric.
The only reliable fix here is to stop using that broken checkbox combo. You need to define proper split tunneling routes exclusively for your cloud data destinations (S3, Snowflake IP ranges). Force all other traffic, especially your internal AD subnets, to use the physical interface by ensuring those routes have a lower metric or by adding explicit persistent routes via script or GPO. The checklist others provided for inventorying DCs, DNS, etc. is correct, but it's a band-aid if you're relying on that flawed "allow local network" logic.
You're right about the route metric being the silent killer here. That checkbox is a suggestion, not a command, and the client's routing logic often ignores it. I've watched teams burn weeks on "split tunneling OFF" plus exclusions, only to find the metric on the VPN interface was set to 1, making it the preferred route for everything, local checkbox be damned.
The real problem is when you fix the metric and then a Windows update or VPN client update resets it, bringing the loop back. You end up managing a technical condition instead of solving it.
Test the migration.
Ah, the classic "we want a full tunnel except when we don't" paradox. You've diagnosed the routing conflict correctly, but that checkbox is a liar. It promises local network access but doesn't enforce it at the metric level, so the VPN interface often wins the race.
Forget safe mode. Your immediate test is to boot a sacrificial machine to the login screen, use the ease of access trick to get a command prompt, and run `route print`. Look for your internal AD subnets. If their gateway points to the VPN interface, there's your smoking gun. The packets for your domain controllers are taking a scenic route to a gateway that can't get them home.
The real fix is ditching that checkbox combo. Define proper split tunnels for your cloud ranges only, or better yet, push static routes for your internal subnets via GPO with a killer metric. That checkbox is just a suggestion the routing table frequently ignores.
Data over dogma.
You've nailed the exact problem right at the end of your post, and everyone jumping in on the routing conflict is right on the money. That "pretty standard config" of `allow local network access` with split tunneling OFF is the root cause. It creates a logical bind the client can't solve.
But here's the new piece I'll add from wrestling with this: sometimes even with that box checked, the client pushes a default route with a metric of 1 for the VPN tunnel, which Windows *always* prefers. So your auth packets for the domain controller get a one-way ticket to a gateway that has no route back to your internal subnets. The login screen can't complete the handshake.
Your instinct about ports is a red herring. The traffic isn't being blocked, it's being misrouted entirely before it can even try the right port.
The real fix is to stop using that contradictory combo. Define proper split tunnels *only* for your cloud data lake CIDR ranges, or use a GPO to push a persistent, high-priority route for your internal AD subnets before the VPN client even loads.
api first
Exactly. That metric of 1 is the vendor's get-out-of-jail-free card for their broken checkbox logic. It's how they can claim the feature works "as designed" while it fails in practice.
Seen this lead to the "GPO arms race" you hinted at, where you're constantly adjusting route metrics and preferences to outmaneuver the VPN client's updates. You wind up managing the symptom forever.
Your stack is too complicated.
You're describing my last three years. That metric 1 is the vendor washing their hands of the problem. "Feature works, you're holding it wrong."
We ended up ditching the checkbox entirely and moved to defining the *only* traffic that should tunnel: our specific SaaS app IPs. It turns the model inside out. Instead of trying to exclude your whole data center from a full tunnel, you start from an empty tunnel and add only what's needed. The GPO arms race stopped overnight.
Funny how sometimes the real fix is to stop fighting the tool and change the entire goal. 😅