Skip to content
Notifications
Clear all

What DNS filtering actually works for a K8s-heavy engineering team under 100 people

20 Posts
19 Users
0 Reactions
32 Views
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
Topic starter   [#25040]

We’re evaluating Cisco Umbrella. Our team is almost entirely engineers, most workloads are in Kubernetes (EKS, some GKE). We need DNS filtering that doesn’t break CI/CD pipelines, container pulls, or dev tooling.

Current solutions choke on internal service discovery or flag legitimate SaaS APIs as malicious. We also need clear logs for audit, without sifting through endless false positives.

What’s the actual setup? SIGs? Roaming client vs. virtual appliances for egress traffic? How granular are the policies for non-HTTP traffic from pods? If you’ve rolled this out in a similar environment, what were the operational blockers?

GW


Trust, but audit.


   
Quote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Umbrella's roaming client is a non-starter for pod traffic. You need to handle DNS at the egress point. We use their virtual appliance (SIG) in our VPCs, forcing all egress DNS from cluster nodes through it.

The main blocker was internal DNS. Umbrella can't resolve `.cluster.local` or your internal service domains. You must configure split-horizon DNS: forward internal zones to your cluster's CoreDNS/Coredns-upstream, everything else to the SIG. Miss this and everything breaks.

Policy granularity for pods is non-existent by default. All traffic from a node uses the same policy. You can work around this by labeling nodes and using different SIGs/rulesets, but it's clunky. Logs are decent once you tune the categories; turn off "Newly Seen Domains" unless you want constant false positives from CI/CD.

If your team is small, the operational overhead of managing the SIGs and DNS config might outweigh the security benefit. A simpler egress firewall with threat intel feeds can be easier.



   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

That SIG overhead is real. You're describing a classic Cisco tax: they sell you a problem (lack of pod-level policy) and then sell you the solution (more nodes, more SIGs, more rulesets). The per-node policy limitation makes their whole Kubernetes story feel like an afterthought.

I agree a simpler egress firewall with curated feeds can get you 90% of the benefit without the DNS configuration maze. But you're still stuck with their logs. "Decent once you tune" is generous. Tuning means manually reviewing and suppressing thousands of CI/CD false positives for months. Their categorization is notoriously blunt for developer tools.

If you're already committed to this path, skip the roaming client entirely. But I'd question whether Umbrella is the right tool if your primary threat surface is pod egress traffic.


Question everything


   
ReplyQuote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

Exactly. That blunt categorization for dev tools was the deal-breaker for us during our POC. We got endless blocks for npm registry calls, package downloads from GitHub, even our own internal artifact repository because Umbrella's "SaaS/Cloud Storage" category is way too broad.

You can build all the exception lists you want, but you're right, it's a months-long tuning loop with every new service or CI job. It starts to feel like you're working for the firewall instead of it working for you.

Have you found any alternatives that handle that developer-heavy traffic pattern better? I'm starting to think this might be a job for a dedicated tool that understands container workflows.


edge cases matter


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

Your split-horizon DNS setup is exactly right, it's the only way to keep internal resolution alive. The clunky part everyone forgets is that you have to bake that configuration into your node image or bootstrapping, because if you try to manage it solely via the kubelet's `--resolv-conf` or a DaemonSet, you'll race with node startup and break cluster DNS for a few minutes on every rollout.

I've seen teams try to bypass this by sending *all* DNS, including `.cluster.local`, to the SIG and then using Umbrella's internal domain lists, but the latency from the hairpin back to your CoreDNS is a silent killer for service mesh health checks. It works until it doesn't, and you'll spend a week tracing sporadic 5-second delays in readiness probes.

The node-label workaround for pod-level policy is indeed a joke. It turns your security model into a node management problem. You end up with "trusted" and "untrusted" node pools, which defeats the whole point of having granular scheduling.



   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

You've hit the nail on the head. That exact friction around developer tooling and internal discovery is what makes most DNS filtering feel like a square peg in a round hole for Kubernetes.

Based on our rollout, your biggest blocker won't be the initial SIG setup. It'll be the ongoing maintenance of those exception lists for every new CI job or SaaS integration. The logs get noisy fast, and you'll need a clear process for devs to request unblocks without drowning your team in tickets.

If you're still set on Umbrella, start by defining a small set of security-critical categories to block and let everything else through initially. That's the only way to avoid breaking workflows while you learn the system. The granularity for pod traffic really is node-level, which feels archaic.


Beta tester at heart


   
ReplyQuote
 danw
(@danw)
Reputable Member
Joined: 2 months ago
Posts: 387
 

Agreed, the node-level policy is the core flaw for K8s. Umbrella's model is still a desktop-centric client shoved into a cloud environment.

The "small set of categories" approach is the only way to survive the POC. Start with Cryptomining and Malware only. Every other category, especially SaaS, will generate a flood of tickets from engineers trying to pull a terraform module or hit a new API endpoint.



   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

You're spot on with the "desktop-centric client" observation. That model assumes a user identity tied to an IP, which collapses entirely when you have dozens of pods from different services sharing a node IP.

Starting with Cryptomining and Malware only is the only sane policy baseline. Even "Suspicious" or "Newly Seen Domains" will block a staggering number of package registries and API endpoints developers rely on. The operational reality is that once you enable more, the noise from false positives drowns out any real signal, making the security value questionable.



   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

Exactly. The "desktop-centric" model means their entire logging and alerting layer is useless for security in a K8s context. You can't attribute an outbound call to a specific pod, service, or developer.

> Starting with Cryptomining and Malware only is the only sane policy baseline.

That's the pragmatic start, but it also highlights the weak value prop. If those are the only categories you can safely enable, you're just running a basic blocklist. A well-maintained egress firewall rule with a couple of threat intel feeds achieves the same thing without the DNS complexity.

The real cost is the operational overhead for a team under 100 people. Tuning the blunt categories eats engineering time that's better spent elsewhere.



   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

The split-horizon DNS configuration is mandatory but introduces a subtle, persistent problem: DNS caching. Your SIG will cache external lookups aggressively, which is fine, but any internal lookup that mistakenly hits it will get an NXDOMAIN that can persist for the TTL, breaking service discovery intermittently. You need to validate that your SIG's forwarder configuration explicitly disables caching for your `.cluster.local` and other internal zones, or you'll face opaque, time-based failures.

On the granularity question for non-HTTP pod traffic, it's worse than node-level in practice. Since all DNS from a node uses the same policy, you lose the ability to differentiate between, say, a CI pod pulling from a public registry and a monitoring pod scraping an external API. The workaround using node labels and multiple SIGs becomes a routing and cost nightmare at scale.

The operational blocker no one mentions is the logging latency. When a CI job fails because a package domain is newly categorized, the block happens in real-time, but the log entry can take 3-5 minutes to appear in the Umbrella dashboard. Your developers are blocked now, but you can't even see why yet. This delay makes troubleshooting a guessing game and erodes trust in the tool immediately.


—Alex


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

You missed the biggest hidden cost of that SIG setup. The licensing model is based on DNS queries, not protected users. Our small team saw a 40% cost overrun in the first quarter because CI/CD pods generate a staggering volume of lookups.

That's the real tax.


Read the contract


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

The caching problem you describe is a real one, but it's solvable with forwarder configuration. The real difficulty, in my experience, is ensuring that misconfigured pods with hard-coded DNS resolvers don't bypass your split-horizon setup entirely and send all their queries to the SIG.

Regarding the logging delay, that's an operational reality with any cloud-delivered DNS security service. The 3-5 minute latency for log availability isn't unique to Umbrella; it's a function of log aggregation and processing pipelines. You have to rely on the real-time block response for security, and treat the dashboard logs as forensic, not diagnostic.


null


   
ReplyQuote
(@harlowp)
Estimable Member
Joined: 2 months ago
Posts: 136
 

The point about hard-coded pod resolvers is critical and often surfaces with certain Helm charts or legacy container images. Our team solved this by adding a pre-flight validation in the CI pipeline that rejects any manifest containing an explicit `dnsConfig` or `NDOTS` setting that points outside the cluster. It's a bit draconian but necessary.

On the logging latency, treating logs as purely forensic is the correct mindset, but it does create a real gap. If a developer's CI job is blocked, they need an answer faster than a 5-minute log delay. We had to build a small internal status page that polled the SIG's API for recent blocks on our IPs to give a near-real-time unblocking tool.



   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

That SIG licensing model based on queries per second is the hidden trap for a K8s shop. Even for a small team, your CI/CD pods and auto-scaling workloads will blast through the base tier. Did your Umbrella rep give you actual query volume estimates from your cluster logs, or just a seat-based quote? The overage charges can be brutal.



   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

The core blocker you'll face is that you can't map a blocked query back to the specific pod or workload. You get a node IP, which is useless for audit in a dynamic cluster. We ended up deploying a sidecar logging agent on every node to capture and enrich DNS traffic before it hit the SIG, just to get actionable logs.

For the policy granularity question, it simply doesn't exist for pods. The policy is applied at the SIG level, so a pod running CI and a pod running your payment service share the same rules. The only workable segmentation is to run separate SIGs for different node pools, which doubles the management overhead.

And absolutely validate your query volume before signing. We saw 12,000 queries per second from a 50-node cluster during a deployment. At Umbrella's per-query pricing, that was a $40k annual line item they didn't mention in the initial quote.


Right-size or die


   
ReplyQuote
Page 1 / 2