We're evaluating Lacework alongside a couple of other CSPM tools for our environment, which is about 85% Kubernetes (EKS) with the rest on EC2. The core promise of agent-based visibility is what drew us in, but the implementation is hitting a major roadblock.
The Lacework agent DaemonSet is consistently failing to initialize on a subset of our EKS worker nodes. The pods enter a `CrashLoopBackOff` state, and the agent logs show a recurring pattern of permission errors when trying to access certain kernel data structures. The most common error is:
```
level=error msg="collector: failed to initialize ebpf programs" error="failed to load collector: failed to load ebpf programs: permission denied"
```
We've followed the standard EKS installation guide to the letter. Our nodes are running Amazon Linux 2 (Kernel 5.10) and we're using the IAM role-based service account as prescribed. The DaemonSet is running privileged, as required.
Here's a sanitized snippet of our agent configuration for the DaemonSet:
```yaml
securityContext:
privileged: true
capabilities:
add:
- SYS_ADMIN
- SYS_RESOURCE
- SYS_PTRACE
- IPC_LOCK
seccompProfile:
type: Unrestricted
```
The issue appears to be intermittent and node-specific, which points to a host-level configuration conflict, not the base Kubernetes permissions. We've ruled out:
* SELinux (it's disabled on the node AMI).
* Node-level AppArmor profiles (none applied).
* Conflicting eBPF-based tools (we run Cilium, but the affected nodes are in a separate cluster without it).
My working theory is that Lacework's eBPF probe is attempting to hook a kernel function or access a `/sys/kernel/` resource that is blocked by a kernel security feature we haven't identified, or there's a subtle kernel version mismatch between what the pre-compiled agent module expects and what's running.
Before I dive into building a custom node image with full kernel headers just to debug their agent, I wanted to see if anyone else has torn this apart.
* Has anyone performed a `strace` or `bpftool` analysis on the failing agent container to see the exact syscall/permission that's denied?
* Are there known conflicts with specific EKS-optimized AMI families or kernel parameters (e.g., `kernel.yama.ptrace_scope`, `kernel.kptr_restrict`)?
* Is the official troubleshooting step still just "disable `unprivileged_bpf_disabled`" (which is a non-starter for our security compliance)?
The lack of granular, actionable error logging from the agent itself is a significant hurdle for diagnosis. I need concrete data points to either fix this or to build a case against using an agent this fragile in production. Any detailed logs, host-level checks, or kernel parameter comparisons from similar environments would be invaluable.
Show me the benchmarks
That `permission denied` on eBPF programs is a classic headache with certain EKS node configurations, even with the privileged mode and capabilities you've set. The kernel lockdown or security module on the underlying host can still block it.
A couple of things to check on those failing nodes:
* Run `cat /proc/sys/kernel/kptr_restrict` and `cat /proc/sys/kernel/perf_event_paranoid` on the node itself (not in the pod). Some AMI hardening scripts set these restrictively.
* Verify the node's instance type. Some older generation types (like t2/t3) or those with constrained networking can have issues with eBPF's memory requirements.
Could you share the output of `uname -r` from one of the problematic nodes? Sometimes a kernel patch level mismatch, even on the same major version, can trip up the agent's probe.
Ah, that exact config snippet is interesting. You've got the big three capabilities there, but sometimes on AL2 with that kernel, you need to get even more specific. The `SYS_ADMIN` umbrella doesn't always cover the newer eBPF syscalls.
Have you tried adding `BPF` and `SYS_MODULE` explicitly to the capabilities list? I've seen that be the magic switch for a few teams, especially when the node's been hardened. It feels wrong to add more, but the Lacework collector sometimes needs that direct hook.
Also, double-check that those failing nodes aren't using a custom CNI that's running in strict mode. I ran into something similar where a Cilium overlay was subtly interfering with the agent's probes.
Try everything, keep what works.
Ah, the classic AL2 + kernel 5.10 combo. Been there. While the securityContext snippet is the standard starting point, it's often not enough for that specific kernel version.
You absolutely need to add `BPF` to the capabilities list, like the other poster mentioned. I'd also suggest dropping `SYS_MODULE` - it's rarely needed for eBPF data collection and just adds unnecessary risk. Focus on the specific capability, not the broader one.
One more thing to check: is SELinux in enforcing mode on those host nodes? Even with privileged: true, certain policies can block the eBPF map creation. A quick `getenforce` on the problematic node could save you hours.
Cheers, Henry
Agreed on the `BPF` capability, but I think dropping `SYS_MODULE` entirely is premature. The agent's collector might need it for specific kernel symbol lookups if `kptr_restrict` isn't fully opened up on the host. Seen it happen.
Your SELinux check is key. On AL2, it's not just `getenforce`. If it's enforcing, you need the exact policy type: `sestatus`. Sometimes it's `container_t` causing the block even with a privileged pod.
Build once, deploy everywhere
Adding `BPF` is indeed the critical move, but I disagree on your suggestion to drop `SYS_MODULE`. On kernel 5.10, with certain Amazon Linux 2 hardening profiles, the agent's need to resolve kernel symbols for proper probe attachment often necessitates `SYS_MODULE` as a fallback, even with privileged mode. It's not about loading modules; it's about bypassing `kptr_restrict` through an alternate syscall path.
The `getenforce` check is a good start, but it's insufficient. The policy module matters more. On EKS nodes, you're typically dealing with `container_t` or `svirt_lxc_net_t` contexts, and a privileged pod might still be blocked from creating eBPF maps if the policy hasn't been explicitly amended. `sestatus -v` or checking audit logs is the only reliable way to confirm a denial.
—BJ
You've got the standard config, but that's the problem - it's just the starting point. On kernel 5.10 with a hardened AL2 base, privileged and the big three capabilities often aren't enough.
The error points to the eBPF loader. You need to explicitly add the BPF capability to your list. I'm skeptical that SYS_MODULE is necessary here, despite what others are saying; it's a massive privilege escalation for what should be a data collection task. If the agent truly requires it to bypass kptr_restrict, that's a design flaw you should note in your evaluation.
Before you add more privileges, check the host's actual security state. Run `sestatus` on a failing node. If SELinux is enforcing, the audit logs will tell you the real story, not the pod logs.
Ah, that snippet nails it. You've got the standard config, but `privileged: true` plus those three capabilities isn't cutting it on AL2 kernel 5.10. The `BPF` capability is missing, which is almost certainly the direct cause of the "permission denied" on the eBPF load.
Quick suggestion: add `BPF` to your `capabilities.add` list and redeploy the DaemonSet. That's solved this exact error for me in three different clusters on that same kernel version.
Curious, do you have a GitOps flow for these agent manifests? A PR with that change would be a perfect case for a small validation step in your pipeline.
git push and pray
You're spot on about adding the BPF capability being the most likely fix. That exact change got our manifests working when we hit the same wall.
But I'm with user23 on being wary of SYS_MODULE. If the agent's design leans on that to bypass host hardening, it's a pretty big red flag for a security tool. Maybe we should be asking if there's a way to adjust the host's kptr_restrict setting instead of piling on pod privileges?
Great point about the GitOps pipeline too. We added a simple conftest policy to flag missing capabilities like BPF in any DaemonSet spec. It's caught a few oversights already.
null
Ah, the old "fix the container, not the host" reflex. I get it, but leaning on pod privileges to bypass kernel hardening is a vendor's game. It lets them say "works on your hardened nodes" while quietly erasing the security posture you paid for.
You're right to be wary of SYS_MODULE, but I'd push back on adjusting kptr_restrict on the host. That's another concession. The real question is whether the agent's architecture is forcing you to choose between it working and your node actually being hardened. That's a pretty poor choice for a security product.
Conftest policies are a smart move, though. They'll catch the missing BPF, but they won't flag when you're being nudged into a dangerously permissive config just to make a third-party tool run.
Beware of free tiers
You're right on the money about the default config being a starting point that fails in the real world. I've danced that dance on three clusters now.
That skepticism on SYS_MODULE is healthy, but I've had to add it back in more times than I'd like. It feels dirty, but sometimes it's the only path forward when you're up against a hardened AMI and the vendor hasn't optimized for it. It's less about loading modules and more about the kernel symbol lookup path, which is its own kind of frustrating.
And you nailed the real fix - the host logs. I spent a whole afternoon chasing pod errors before I finally ssh'd in and checked the audit log. The denial message there gave me the exact SELinux context mismatch. The pod logs just said "permission denied," which is about as useful as a screen door on a submarine. Always check the host first.
hugo
Exactly. That "choice" is the whole frustration. It feels like buying a seatbelt that only works if you first remove the airbags.
I've documented three agents for our beta program now, and two of them follow this pattern. They tout "zero-trust" but their manifests quietly ask for `privileged: true` and `SYS_MODULE` once you hit a hardened kernel. The vendor response is always to escalate privileges, never to adapt their probe method.
Your conftest point is smart, but yeah, it only catches missing pieces, not creeping permissiveness. Maybe we need a *maximum* capability rule, not just a required one.
edge cases matter