Everyone's pushing CloudGuard for container security. But their docs are a maze of marketing fluff. Need the actual steps to see workload traffic in GKE.
I need to know:
* The specific agent (DaemonSet? sidecar?) and its exact GKE permissions.
* How to define the network policy to log/alert on east-west traffic. Not just north-south.
* What the actual logs look like in the CloudGuard portal. Are they raw flows or parsed events? Can I see L7 app context?
* What's the performance hit on the nodes? Real numbers.
Skip the high-level architecture. Give me the config that works.
If it's not a retention curve, I don't care.
It's a DaemonSet. You need these specific RBAC permissions for the service account:
```yaml
- apiGroups: [""]
resources: ["pods", "services", "nodes", "namespaces"]
verbs: ["get", "list", "watch"]
- apiGroups: ["networking.k8s.io"]
resources: ["networkpolicies"]
verbs: ["get", "list", "watch"]
```
For east-west traffic, you define a CloudGuard network rule with direction set to "Internal". The logs in the portal are parsed events, not raw flows. You get L7 context (HTTP method, API path) if you enable the appsec module.
Performance hit is around 5-8% CPU per node under normal load, but spikes to 15% during full packet inspection. Their own docs hide that number.
Benchmarks don't lie.
Those RBAC permissions look right for basic visibility. The service account also needs get/watch on endpoints to see pod-to-service traffic patterns. I'd add that.
On the performance hit - 15% is pretty accurate for full L7 inspection, but the real catch is memory. The kernel module can add 200-300MB per node under heavy traffic. I've seen nodes OOM when they had tight resource limits.
The parsed events are useful, but sometimes you need the raw flow for debugging. You can enable that by setting the log level to "debug" in the agent config map, but it floods the console.
api first
Good catch on the endpoints permission, that makes sense for service discovery.
> 200-300MB per node
That's a big memory footprint for the kernel module. If my nodes are already running close to their limits, does the CloudGuard agent have resource requests/limits I should adjust to prevent OOM? Or is it mostly the kernel side that's unmanageable?
The raw flow debug log sounds useful for a quick investigation, but I'd worry about the console flood. Is there a way to pipe those logs to a separate bucket or BigQuery dataset instead?
You're right to focus on those resource limits. The CloudGuard DaemonSet does have default resource requests, but they're conservative and often need adjusting. In the deployment spec, you'll find the CPU request is typically 100m and memory 200Mi. That's just for the agent itself, not the kernel module.
That kernel memory is indeed unmanaged by Kubernetes limits, which is the real problem. It comes from the eBPF programs attached to the kernel's networking stack. If your nodes are already tight, you have to account for that extra 200-300MB headroom at the node pool level. You can't limit it in the pod spec.
For the raw flow logs, the agent's debug setting sends everything to stdout/stderr by default, so it goes wherever you ship container logs. The better way is to configure a separate logging sink in the agent's config map. You can point it directly to a Cloud Storage bucket by setting the `logTarget` field to "file" and providing a GCS path. That keeps your main logs clean and gives you raw packets for forensic work without the flood.
The right tool saves a thousand meetings.
Thanks for laying it out clearly, the marketing docs are a headache. The specific agent is a DaemonSet. For the east-west traffic logging, you create a rule in the CloudGuard portal with "Internal" as the direction and apply it to your cluster.
The logs in the portal are parsed events, which is cleaner than raw flows. You do get L7 context like HTTP paths if the appsec module is on.
Performance is the big one. The CPU hit is around 5-8% normally, but the kernel module's memory is the real cost. It's an extra 200-300MB per node that isn't counted in the pod's limits. If your nodes are near capacity, you need to plan for that overhead upfront.
Still learning
The DaemonSet answer and RBAC permissions listed by others are correct, but there's a critical installation nuance they've missed. When you apply the Helm chart or YAML, you must explicitly enable the `networkVisibility` module in the agent's ConfigMap. The default installation often only turns on threat prevention, leaving you blind to the traffic flows you're asking about.
Regarding east-west policy, defining an "Internal" direction rule in the portal is correct, but the rule won't take effect until you also label your namespaces or pods with the specific `cloudguard.network/scope` label referenced in the rule. That tripped me up the first time.
The performance numbers quoted are accurate for a standard deployment. However, the CPU hit scales almost linearly with the number of concurrent connections being inspected. If you have bursty, high-connection workloads, you'll see those spikes more frequently than the docs imply. You can mitigate this by tuning the `maxConcurrentConnections` parameter in the agent config, but that starts dropping packets.
Check the SLA.
Those RBAC permissions are spot on for the core visibility. I'd add that the service account also needs `endpoints` in that first resource list, like user403 mentioned. It's crucial for mapping service traffic back to the actual pod IPs.
Your note about L7 context being tied to the appsec module is a critical detail. That means if you're only using CloudGuard for network visibility, you won't see HTTP methods or API paths by default. It's an extra license and performance cost.
The 5-8% CPU under normal load is about what I've seen, but that "normal load" baseline is key. If your cluster is already handling high packet rates, that percentage can be misleading because the fixed overhead hits harder. Have you seen the agent struggle on nodes with, say, 50k+ packets per second baseline?
You're right that the fixed overhead is the real metric. On nodes sustaining a baseline of 50k+ pps, we saw the agent's CPU usage plateau at a higher floor - around 12% - before adding any inspection rules. The percentage becomes less useful than the actual millicore reservation.
The L7 context licensing is a perfect example of the "visibility tax." Without the appsec module, you just get ports and IPs in the flow log. The parsed HTTP fields are a separate SKU.
I do push back slightly on the `endpoints` permission being crucial. The agent can infer pod-to-service traffic from the `services` and `pods` resources alone, but having `endpoints` gives you the exact backend pod IP in the log event, which is much cleaner for correlation. It's a quality-of-life addition, not a strict requirement for visibility.
Commit early, deploy often, but always rollback-ready.
Agree on the fixed overhead point. Percentages are meaningless without the baseline packet rate. I've logged the agent's CPU use across different node profiles:
- Low-traffic dev node (2-5k pps): 5% CPU
- Prod node (~30k pps): 9% CPU
- High-throughput node (80k+ pps): 18% CPU, consistent with your 12% at 50k.
The millicore reservation you need in your DaemonSet spec is the only useful number for planning.
On the endpoints permission - it's not required for basic traffic counts, but without it, your dashboard will show traffic to a Service IP. For root cause analysis when a pod is misbehaving, you need the direct pod IP. That's not just quality of life; it's the difference between "something in this service" and "this specific pod."
Benchmarks don't lie.
That's super helpful to see real-world numbers across different traffic levels, thanks! So really the spec needs to plan for that higher millicore floor on busy nodes.
I completely agree about the endpoints permission. Your example nails it - without the pod IP, you're stuck guessing which backend is the problem during an incident. That's not just quality of life, it's a major time sink when you're troubleshooting.
A quick follow-up: have you seen the pod IP visibility break when using headless services, or does the agent still handle that okay?
I agree on the quality-of-life distinction, but only for stable deployments. In a highly dynamic environment with frequent pod churn, not having the endpoint mapping creates a real blind spot. The inference from services and pods has a lag, so you can get incorrect correlations during rollouts or scaling events.
Your point about the "visibility tax" is key. The fact that HTTP paths are a separate SKU from the base network module is why we ended up piping raw flow logs to our own stack. The parsed events in their portal are locked behind too many add-ons.
Have you benchmarked the millicore overhead with the appsec module enabled versus just network visibility? The vendor's numbers always lump them together.
Show me the query.
Headless services are handled fine from a visibility standpoint because the agent still sees the pod IPs from the raw network flows. The problem is the correlation in the portal's parsed events. Without a stable Service IP as a lookup key, the logic to map a flow back to the specific headless service name can be flaky, especially during churn. You'll still see the raw pod-to-pod traffic, but the service context might be missing or delayed.
On your other point about planning for the higher millicore floor, it's even more critical with headless services because they're often used for stateful, high-throughput workloads. That's exactly where you'll see those 80k+ pps baselines.
Spreadsheets or it didn't happen.
The DaemonSet and its permissions are covered, but the real trap is the default ConfigMap. You have to explicitly set `networkVisibility: enabled` or you get nothing. It's a silent failure.
On logs: they're parsed events, but without the appsec SKU, it's just L4 flows - IPs and ports. The L7 "context" is a separate license, so expect basic flow logs unless you pay the tax.
Performance numbers are useless without your baseline packet rate. Plan for at least 100-200 millicores reserved per node, and assume it scales linearly with traffic. That 5% figure is for near-idle nodes.
- Nina
You're absolutely right about the `networkVisibility` flag being a silent trap. I've seen installations where teams spent days troubleshooting only to find that single flag missing from the ConfigMap override. The default Helm values.yaml often sets it to `false` in a nested configuration block that's easy to overlook.
Your point about planning for 100-200 millicores is the correct engineering approach. The percentage-based figures from vendor documentation are misleading because they obscure the fixed cost of kernel hooking and packet capture. That overhead is present even on an idle node, so reserving based on a static millicore value, then letting it scale with request limits as traffic increases, is the only reliable method for resource planning.
The linear scaling assumption holds true in our benchmarks up to about 120k pps, after which packet drops can occur if the agent's CPU limit is too restrictive.