Hey everyone! Ran into a tricky one with our FortiSASE setup this week and hoping someone has already mapped this out.
Our team uses cloud-based CI/CD runners (think GitLab, but also applies to others) to build and deploy. Ever since we rolled out FortiSASE with full tunnel client traffic inspection, those jobs have been failing with random timeouts. The logs point to network latency and packet loss during large dependency downloads (like pulling docker images or npm packages). It's clearly the inspection causing the bottleneck for these high-throughput, automated connections.
I know we can create policies to exclude specific IPs or services from inspection, but I'm a bit fuzzy on the *best practice* way to do this for something like a dynamic cloud service. The runners' source IPs can change.
What we're considering:
* Creating a FQDN address object for the cloud runner service domain and exempting it.
* Setting up a separate, non-inspected policy based on a tags or user groups (maybe a "service-account" user).
* Is there a way to reliably identify this traffic by application signature instead?
Has anyone successfully built an exclusion for their CI/CD infrastructure without opening a huge security gap? I'd love to hear how you structured your policy and what objects you used. A step-by-step of your rule order would be super helpful!
Keep it simple.
Oh man, this is a classic SASE growing pain. We hit the exact same thing with our Azure DevOps agents last year.
Your FQDN object idea is probably the most straightforward path. That's what ended up working for us. Just make sure you're catching the actual domain the runners use to pull dependencies, which might be different from their management API domain. We had to whitelist a couple of CDN domains for package managers too.
A quick heads-up: if you go the user group or tag route for a service account, double-check how the traffic sourcing works. Sometimes the tunnel client on a VM won't pass the identity context the same way a user device does, and the policy might not match.
Yep, the identity context point is critical. We had a policy based on AD groups that flat out didn't trigger for our Jenkins nodes because the tunnel client wasn't doing user-based auth.
You'll need to match on source IP or a tag applied directly to the VM. We used a static tag based on the cloud instance role.
Data over opinions
Exactly. Source IP is the only consistent identifier for a runner VM in this setup. Even a tag can be tricky if your IaC doesn't apply it before the tunnel client starts.
Your best bet is to carve out a dedicated subnet for the runners and create a policy based on that CIDR block. Static, reliable, and easy to audit.
I generally agree that a source IP CIDR block is the most static and foolproof identifier, but I've found this approach can create a significant security blind spot over time. That dedicated subnet becomes a privileged network path that bypasses all threat inspection, which is fine for your own package mirrors but risky if a runner ever gets compromised and needs to pull from an external, malicious repository. The policy becomes a blanket exemption for all traffic from that IP range.
A slightly more nuanced alternative we implemented is to combine the CIDR-based policy with specific FQDN objects for the allowed service endpoints (like pkg.jenkins.io, registry-1.docker.io, artifacts.maven.org). This maintains the reliability of IP matching for the source but still inspects traffic leaving that subnet for destinations not on the pre-approved list. It adds a bit more policy management overhead, but it limits the trust boundary.
Support is a product, not a department.
That's a really interesting hybrid approach. You get the reliability of the source IP match, but you're right, it cuts down the risk surface a lot compared to a full bypass.
A follow-up question on that, though. How do you handle it when one of those services, like a package registry, uses a huge pool of CDN hostnames that can change? We had to whitelist a domain once, only to find the actual downloads came from a different, rotating set of addresses. Do you just keep expanding the FQDN object list, or is there a way to handle that gracefully?