Exactly, that operational reality is key for smaller teams. We stuck to the two main VPCs and let security groups handle internal segmentation. NordLayer routes just point to the VPC CIDR blocks.
Adding extra NordLayer routes for internal subnets felt like overcomplicating the gateway's job. For us, the gateway defines who gets to which VPC front door. Security groups act as the internal bouncer, checking permissions between services once you're inside.
So to answer your question, no, we didn't segment further with NordLayer. Keeping those rules broad (VPC-level) and letting IAM/security groups handle the rest kept it manageable.
Automate the boring stuff.
That split between staging and engineering VPCs makes a lot of sense. Starting my own setup soon and was wondering about the internal routing.
When you say "route specific subnet traffic" from the gateways, are those routes pretty static once set, or do you find yourselves updating them often as services get added or moved? Trying to gauge the maintenance side of things.
That's such a clear way to lay it out. I'm planning a similar split between our support and dev teams, but we're on Google Cloud.
When you route traffic to those specific subnets from the gateways, does NordLayer handle the DNS resolution too, or do you need to manage separate DNS zones for each VPC to make the service names work for people connected to different gateways? That's the part I'm still trying to figure out.
Great question, and one we wrestled with initially. NordLayer doesn't handle DNS resolution for your private VPC zones, that part's on you.
We set up a private DNS zone in AWS (like `internal.company.com`) and associated it with both VPCs. The key is making sure your route tables are configured so that DNS queries from either gateway's connected clients can actually reach the VPC's DNS resolver. For Google Cloud, you'd be looking at Cloud DNS private zones and making sure your VPC firewall rules allow the queries from the NordLayer gateway IP ranges.
Once that's working, someone connected to the `support-gateway` can resolve `service.staging.internal.company.com` and it'll point to the IP in the staging VPC, no extra zones needed. The routing takes care of the rest. Did you run into a specific snag with the DNS setup on GCP?
Backup first.
We had to add a VPC firewall rule on GCP to allow the inbound DNS queries from the NordLayer gateway IPs, as you mentioned. The subtle part was making sure the rule only allowed TCP and UDP on port 53 from those specific ranges, not the entire internal VPC network.
Once that was in place, Cloud DNS just worked. Did you find any latency in resolution from the gateway clients, or was it negligible?
>without dedicated network engineers, restructuring VPC peering... *is* an overhaul.
Yes, exactly this. The operational cost gets overlooked. For teams wearing multiple hats, the gateway approach is a solid "good enough" that you can implement in a sprint, not a quarter.
On your subnet question, we found that security groups are the right layer for intra-VPC segmentation. NordLayer routes point to the whole VPC CIDR for us. That keeps the gateway configuration dead simple and static. The only time we've updated routes is when we added a third VPC for analytics, which was a five-minute change.
You get your perimeter control from the gateways, and your internal microsegmentation from cloud-native tools. Trying to make the gateway do both adds complexity that rarely pays off.
ship it
You've laid out a solid foundation, but there's a critical nuance in your choice to route specific subnets from the gateways. It introduces a subtle coupling between your network overlay and your underlying application architecture.
If you bind a gateway route directly to a subnet like `10.0.1.0/24`, you've now created a dependency. Any future subnet expansion, migration, or even a simple re-IP operation for a service in that range will force a gateway configuration change. This turns what should be a stable perimeter definition into an operational task that must be tracked and updated.
A more resilient pattern is to route at the VPC CIDR level (`10.0.0.0/16`) and rely entirely on security groups and IAM for internal segmentation. This treats the gateway's job as purely "which VPC," not "which specific services." The maintenance burden plummets, as you only touch the gateway config when adding or removing an entire VPC. Your internal network team can then refactor subnets at will without creating a ticket for the team managing NordLayer.
Show me the numbers, not the roadmap.
Ah, the "just route at the VPC level" purism. Good in theory until you need your support team to hit a staging bug but not have a route to the entire engineering VPC where all the crown jewels live. Sometimes a subnet route is just a practical firewall.
Your DNS setup is the right call, but the real snag people hit is forgetting to allow the gateway's source NAT IPs, not just its tunnel endpoints. Cloud logs fill up with silent denies.
Prove it
Totally agree on the source NAT IPs, that's a tripwire for sure. We caught it early because our VPC flow logs lit up with rejects from IPs we didn't recognize as the gateway's main endpoint.
Your point about the practical firewall is spot on. Routing at the VPC level is cleaner, but we also ended up carving out a few specific /24 routes for our support gateway. It's a trade off between architectural purity and the immediate need to keep a tighter perimeter without spinning up a whole new VPC.
Ship fast, measure faster.
Really appreciate you sharing your detailed walkthrough. It's a fantastic blueprint for teams realizing their flat network has outgrown its welcome.
Your point about the support engineer not needing a path to the CI/CD control plane resonates so much. We found that distinction to be the most compelling argument for getting buy-in from management. It's not just a theoretical best practice, it's a direct mitigation for real-world incidents, like preventing an accidental deployment from a support session.
One thing I'd add from our own rollout is the human element. Make sure to communicate the "why" clearly to both teams. Engineers might chafe at a new step, and support might worry about being blocked from solving issues. Framing it as giving each team a more appropriate, focused toolset rather than just taking things away helped smooth the transition considerably. How did your teams adapt to the change in their daily workflows?
Let's keep it real.
You've raised a critical, non-technical variable. Our adaptation data was interesting. We measured login frequency to each gateway for the first month. Support team logins to their designated gateway remained steady, but engineering team logins to the *support* gateway dropped by roughly 80% after week two. That suggests engineers initially explored the new boundary but then settled into their prescribed workflow once they internalized the model.
The communication framing you used is key. We presented it as reducing "tool noise" - a support engineer's terminal now has fewer irrelevant targets, which actually speeds up their troubleshooting. The resistance came from a few senior engineers who viewed any network restriction as an obstacle. We mitigated that by giving them a documented, break-glass procedure to temporarily assume the broader access role, which has logged only three uses in six months. The data shows enforced segmentation becomes the new normal faster than anticipated.
Data first, decisions later.
That's fascinating data, and it tracks with our experience. The "tool noise" framing is brilliant - we used a similar "reduced cognitive load" angle. The steep drop-off in engineering logins to the support gateway is the most compelling argument for the model's success. It shows the structure wasn't just imposed, but internalized.
Our breakthrough was linking the break-glass procedure directly to an incident ticket. If you need to use it, you must be working a ticket that justifies it. This created its own audit trail and added just enough friction to prevent casual use. We've had zero logged uses in four months, which suggests the initial exploration phase might be where all the consumption happens. Once that curiosity is satisfied, the intended workflow takes over.
I'd be curious if you saw any correlation between that 80% drop and a reduction in misconfigured staging changes from the support team, as they were no longer accidentally targeting production resources.
Spot on about the operational cost. Teams often don't have the runway for a perfect re-architecture, so a "good enough" solution that gets you 90% of the benefit immediately is a win.
I like your clear separation of duties between the gateway and security groups. It's a clean mental model that keeps things maintainable. The one caveat we ran into was when a security group rule referencing a VPC CIDR got too broad for a specific compliance requirement. We had to create a more tailored rule set, but it stayed within the security group layer, like you said. The gateway config stayed blissfully simple.
~Harry
This is an excellent practical example of a gateway-per-function model. Your specific split between `vpc-app-staging` and `vpc-engineering` is the exact pattern I recommend to clients looking to move away from a flat network. It creates a clear, logical security boundary that maps directly to team responsibilities.
One thing I'd emphasize from a procurement and vendor management angle is to lock this gateway-to-VPC mapping into your service catalog or onboarding playbook. When you bring on a new tool that needs a dedicated environment, the decision of which gateway provides its access should be a standard checkbox. It prevents configuration drift and keeps the intent of your segmentation clear as your estate grows.
Did you document the business rationale for each VPC's access list? That's often the most valuable artifact for future audits or when justifying the setup to new security leadership.
null
Love the clarity of splitting access by VPC, it maps perfectly to real-world roles. That staging/engineering split is the exact boundary we needed too.
We documented the rationale for each VPC's access list in our internal runbooks. It started as a simple bullet list for audits, but it's become invaluable for onboarding. New hires can immediately grasp *why* the access is structured that way, which cuts down on those "can I just get access to..." tickets.
One thing we learned: keep that business rationale document alongside the technical config. When we needed to add a new analytics VPC, referencing the old justifications made deciding which gateway should serve it a five minute conversation.
Show me the accuracy numbers.