In our deployment pipelines, we require absolute certainty that our CI/CD servers cannot leak traffic outside our sanctioned networks during sensitive operations, such as pushing artifacts to a private registry or executing security scans. A network-level kill switch is essential for this guarantee. This guide details our implementation of NordLayer as that kill switch on our Jenkins and GitHub Actions runners.
The core principle is to configure the NordLayer client to block all non-VPN traffic and then integrate its control into our pipeline initialization scripts. We run our builders on Ubuntu 22.04 LTS. The following steps were automated via Ansible, but I'll present the shell commands for clarity.
**1. Installation and Systemd Service Configuration**
We install the NordLayer CLI and configure it to start with the kill-switch enabled.
```bash
# Add NordLayer repository and install
curl -s https://downloads.nordlayer.com/linux/latest/deb/nordlayer_linux_amd64.deb -o /tmp/nordlayer.deb
sudo apt install /tmp/nordlayer.deb -y
# Set login credentials via environment variables (secrets managed by our vault)
sudo nordlayer login --credentials $NORDLAYER_USERNAME $NORDLAYER_PASSWORD
# Configure the service to enforce the kill switch
sudo nordlayer settings set KillSwitch true
sudo nordlayer settings set AutoConnect true
sudo systemctl enable nordlayer.service
```
**2. Integration into Pipeline Job Definition**
Our Jenkins `Jenkinsfile` now includes a pre-condition check and a graceful handling mechanism. The pipeline will fail fast if the VPN is not active.
```groovy
pipeline {
agent any
stages {
stage('Validate Network Security') {
steps {
script {
def nordStatus = sh(script: 'sudo nordlayer status --json', returnStdout: true).trim()
def statusJson = readJSON text: nordStatus
if (statusJson.state != 'CONNECTED') {
error("NordLayer kill switch not active. Pipeline halted. VPN state: ${statusJson.state}")
}
echo "✓ Kill switch active. Proceeding from IP: ${statusJson.ip_address}"
}
}
}
stage('Build and Push') {
steps {
// Your secure build and deployment steps here
sh 'docker build -t internal.registry.com/app:${BRANCH_NAME} .'
sh 'docker push internal.registry.com/app:${BRANCH_NAME}'
}
}
}
post {
failure {
// Optional: Notify on network compliance failure
echo "Pipeline failed due to network security violation or other error."
}
}
}
```
**Key Observations & Pitfalls:**
* **Service Account Management:** The NordLayer service runs under `nordlayer:nordlayer`. Ensure your CI user has the necessary `sudo` permissions for `nordlayer` commands without a full password prompt, configured via `/etc/sudoers.d/`.
* **DNS Considerations:** We encountered Docker network issues where containers bypassed the VPN. The solution was to bind Docker to the NordLayer interface (`nordtun0`) by configuring `daemon.json`:
```json
{
"dns": ["10.5.0.1"],
"iptables": false
}
```
* **Connection Stability:** We schedule periodic connectivity checks in a cron job to restart the service if it drops, as a transient network failure on the host could otherwise expose traffic.
This setup has provided the enforced network egress control we required. It adds approximately 1-2 seconds to pipeline startup for the validation check, which is an acceptable trade-off for the security assurance. I am interested if others have implemented similar patterns with other providers or have optimized the failover procedure further.
--crusader
Commit early, deploy often, but always rollback-ready.
This is a smart approach for hardening isolated CI/CD runners. I've seen similar setups where the kill-switch enforcement wasn't verified until a failure occurred, causing a cascade of blocked legitimate traffic. Do you have a health check or monitoring step that confirms the kill-switch is actively blocking before your sensitive pipeline stages run?
Reviews build trust.
Absolute certainty? You're outsourcing that certainty to a third party client. What's in their kill switch, an iptables wrapper? If that client crashes or their daemon hangs, does your 'guarantee' hold?
Your vendor is not your friend.
Third party dependency is my main concern too. The licensing terms matter if this becomes part of your infrastructure. Did you review if commercial use in an automated CI/CD setup violates NordLayer's ToS? I've seen services charge a different tier for server use.
How are you handling the service account cost? Is it a per-server license, and does the price jump at renewal?
Great question. We do run a quick health check. Right after connecting NordLayer, the pipeline pings a known external IP and expects a timeout. If it gets a reply, we fail the build immediately. It's not perfect, but it catches a daemon crash before any real work starts.
You're right about cascade failures - we learned that the hard way when a DNS leak slipped through. 😅 Now we also log the connection state at the start of each sensitive stage.
Docs save time
Good point about the external IP ping test. That's a clever way to verify the kill switch is engaged before any data moves. It catches a total daemon failure.
I'd also consider probing for a DNS leak specifically, since that's where you got burned. You could have the health check script resolve a test hostname and verify the returned IP is within your VPN's range, not your ISP's resolver. It adds one more check against a partial failure state.
That's a really solid addition to the health check. The DNS leak test is crucial because the kill switch might be active on the main interfaces, but if the VPN's DNS isn't properly forced, your queries can still slip out and expose what you're connecting to, even if the data itself routes through the tunnel.
We actually script a similar check using `dig`. We have a dedicated, public test hostname that only resolves to a specific IP if you're using our VPN's DNS servers. If the resolved IP matches our ISP's resolver result, we kill the pipeline. It adds maybe two seconds, but the peace of mind is worth it.
Have you found any edge cases where a DNS proxy or a local cache could give a false positive on that check?
hannah
The dedicated test hostname is a great idea for a controlled DNS leak test. We use a similar approach by querying a known public DNS logging service like `dnsleaktest.com` from within the pipeline, but your method is cleaner and doesn't rely on an external service's availability.
> Have you found any edge cases where a DNS proxy or a local cache could give a false positive on that check?
Yes, local caching can definitely bite you. Our early checks used a static hostname, and we'd get a cached result from a previous, non-VPN run, causing a false pass. We now append a unique subdomain (like `{BUILD_ID}.vpn-test.yourdomain.com`) to the query to guarantee an uncached lookup. Also, some container environments run a local DNS proxy (like the Docker daemon's `127.0.0.11`) which can complicate things if not explicitly configured to forward through the tunnel.
Cloud cost nerd. No, I don't use Reserved Instances.
Good to see the specific steps, especially the systemd integration. That's often where these setups get brittle. I've seen issues where the systemd service starts, but the kill-switch rules aren't fully applied yet, leaving a small window for leaks. Does your pipeline script wait for a specific status from `nordlayer status` before proceeding, or does it just rely on the service being active?
Stay curious, stay critical.
Good call on waiting for the service to be active, but we found that `systemctl is-active` returning "active" doesn't mean the kill switch rules are fully applied. There's a race condition.
Our script now polls for a specific connection status. After starting the service, we wait for `nordlayer status` to output "Connected" *and* we verify the expected virtual interface is up.
Something like:
```bash
while [[ $(nordlayer status 2>/dev/null | grep -c "Connected") -eq 0 ]]; do
sleep 2
done
```
Do you think that's sufficient, or are there other iptables/nftables states we should be checking for too?
Data is the new oil - but it's usually crude.
Checking the interface is the right move, but grep for "Connected" can still be a false positive if the daemon reports connected before the routing table is updated.
I'd also verify the default route points to the tunnel interface. Add a check like `ip route show default | grep -q tun0` (or whatever your interface is) after the status loop.
That confirms the kernel networking stack is actually using the tunnel, not just that the client thinks it's connected.
slow pipelines make me cranky
Absolutely, checking the default route is the critical piece. I've been burned by that exact race condition before - the client says "Connected" but the routing table hasn't converged yet.
We actually combine that route check with a quick `curl` to a known internal endpoint that only answers on the VPN network. The route check confirms the path, and the curl validates we can actually talk to something through it. It adds maybe one more second, but catches those weird states where the route exists but the tunnel isn't passing traffic yet.
> `ip route show default | grep -q tun0`
One caveat: if you're using a split-tunnel config (which some CI setups do for speed), the default route might not point to `tun0`. In that case, you'd need to verify the specific routes for your protected destinations are correctly pointed at the tunnel interface.
Integration Ian
Clever, but you're just adding a second single point of failure. Now your CI depends on your DNS infrastructure *and* NordLayer's VPN. A simple IPv6 leak could still blast right past your clever `dig` test.
Your vendor is not your friend.