Yep, that `networkVisibility` flag has bitten my team too! We wasted a day because we assumed enabling the module in the UI automatically flipped it in the Helm chart. Nope.
> planning for 100-200 millicores is the correct engineering approach
Totally. We started by reserving exactly that, but added a 20% buffer after a node spike during a deployment saturated the agent and caused packet loss. It's that fixed kernel hooking cost - it doesn't care if your traffic is "normal" or a burst.
Your benchmark note about 120k pps is interesting. Was that with a specific CPU limit configuration? We've seen the linear scaling start to bend around 100k if the limits aren't generous enough for the packet processing queue.
Keep deploying!
You're right, the docs are a mess of fluff. Here's what you actually need to deploy.
The agent is a DaemonSet. The exact RBAC you need is a ClusterRole with get, list, watch on `pods`, `services`, and `endpoints`. Don't skip endpoints like some suggest. Without it, you see traffic to a Service IP, not the offending pod IP, which is useless during an incident.
The network policy definition is done in their portal, not via Kubernetes NetworkPolicy. You create a rule with source and destination set to your cluster's internal CIDR, set the action to "Log", and apply it to your cluster asset. That's what captures east-west.
In the portal, without the appsec license, logs are parsed L4 flow events: timestamps, source/dest IPs and ports, protocol, bytes. No HTTP paths or headers. It's a separate SKU.
Performance hit is a fixed overhead plus scaling. On an idle node, reserve at least 100 millicores. On a node sustaining 50k packets per second, it will use 250-300 millicores. The vendor's percentage figures are meaningless. You must set explicit CPU requests and limits in the DaemonSet spec or it will get throttled and drop packets.
Thanks for sharing these exact permissions, that's really helpful for someone just trying to get it running. I appreciate you calling out the performance numbers from the docs too.
The point about logs being parsed events, not raw flows, clarifies a lot about what we'll actually see. So even the basic network logs are processed, just at the L4 level.
One quick follow-up: when you say the CPU spikes to 15% during full inspection, is that triggered by a specific rule, or is it just when traffic volume itself is high?