Skip to content
Notifications
Clear all

Help: Agent keeps failing on our EKS nodes with 'permission denied' errors.

12 Posts
12 Users
0 Reactions
27 Views
(@avag2)
Honorable Member
Joined: 3 months ago
Posts: 376
Topic starter   [#23305]

We're evaluating Lacework alongside a couple of other CSPM tools for our environment, which is about 85% Kubernetes (EKS) with the rest on EC2. The core promise of agent-based visibility is what drew us in, but the implementation is hitting a major roadblock.

The Lacework agent DaemonSet is consistently failing to initialize on a subset of our EKS worker nodes. The pods enter a `CrashLoopBackOff` state, and the agent logs show a recurring pattern of permission errors when trying to access certain kernel data structures. The most common error is:

```
level=error msg="collector: failed to initialize ebpf programs" error="failed to load collector: failed to load ebpf programs: permission denied"
```

We've followed the standard EKS installation guide to the letter. Our nodes are running Amazon Linux 2 (Kernel 5.10) and we're using the IAM role-based service account as prescribed. The DaemonSet is running privileged, as required.

Here's a sanitized snippet of our agent configuration for the DaemonSet:

```yaml
securityContext:
privileged: true
capabilities:
add:
- SYS_ADMIN
- SYS_RESOURCE
- SYS_PTRACE
- IPC_LOCK
seccompProfile:
type: Unrestricted
```

The issue appears to be intermittent and node-specific, which points to a host-level configuration conflict, not the base Kubernetes permissions. We've ruled out:

* SELinux (it's disabled on the node AMI).
* Node-level AppArmor profiles (none applied).
* Conflicting eBPF-based tools (we run Cilium, but the affected nodes are in a separate cluster without it).

My working theory is that Lacework's eBPF probe is attempting to hook a kernel function or access a `/sys/kernel/` resource that is blocked by a kernel security feature we haven't identified, or there's a subtle kernel version mismatch between what the pre-compiled agent module expects and what's running.

Before I dive into building a custom node image with full kernel headers just to debug their agent, I wanted to see if anyone else has torn this apart.

* Has anyone performed a `strace` or `bpftool` analysis on the failing agent container to see the exact syscall/permission that's denied?
* Are there known conflicts with specific EKS-optimized AMI families or kernel parameters (e.g., `kernel.yama.ptrace_scope`, `kernel.kptr_restrict`)?
* Is the official troubleshooting step still just "disable `unprivileged_bpf_disabled`" (which is a non-starter for our security compliance)?

The lack of granular, actionable error logging from the agent itself is a significant hurdle for diagnosis. I need concrete data points to either fix this or to build a case against using an agent this fragile in production. Any detailed logs, host-level checks, or kernel parameter comparisons from similar environments would be invaluable.


Show me the benchmarks


   
Quote
(@integrations_jane_new)
Estimable Member
Joined: 6 months ago
Posts: 155
 

That `permission denied` on eBPF programs is a classic headache with certain EKS node configurations, even with the privileged mode and capabilities you've set. The kernel lockdown or security module on the underlying host can still block it.

A couple of things to check on those failing nodes:
* Run `cat /proc/sys/kernel/kptr_restrict` and `cat /proc/sys/kernel/perf_event_paranoid` on the node itself (not in the pod). Some AMI hardening scripts set these restrictively.
* Verify the node's instance type. Some older generation types (like t2/t3) or those with constrained networking can have issues with eBPF's memory requirements.

Could you share the output of `uname -r` from one of the problematic nodes? Sometimes a kernel patch level mismatch, even on the same major version, can trip up the agent's probe.



   
ReplyQuote
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
 

Ah, that exact config snippet is interesting. You've got the big three capabilities there, but sometimes on AL2 with that kernel, you need to get even more specific. The `SYS_ADMIN` umbrella doesn't always cover the newer eBPF syscalls.

Have you tried adding `BPF` and `SYS_MODULE` explicitly to the capabilities list? I've seen that be the magic switch for a few teams, especially when the node's been hardened. It feels wrong to add more, but the Lacework collector sometimes needs that direct hook.

Also, double-check that those failing nodes aren't using a custom CNI that's running in strict mode. I ran into something similar where a Cilium overlay was subtly interfering with the agent's probes.


Try everything, keep what works.


   
ReplyQuote
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Ah, the classic AL2 + kernel 5.10 combo. Been there. While the securityContext snippet is the standard starting point, it's often not enough for that specific kernel version.

You absolutely need to add `BPF` to the capabilities list, like the other poster mentioned. I'd also suggest dropping `SYS_MODULE` - it's rarely needed for eBPF data collection and just adds unnecessary risk. Focus on the specific capability, not the broader one.

One more thing to check: is SELinux in enforcing mode on those host nodes? Even with privileged: true, certain policies can block the eBPF map creation. A quick `getenforce` on the problematic node could save you hours.


Cheers, Henry


   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

Agreed on the `BPF` capability, but I think dropping `SYS_MODULE` entirely is premature. The agent's collector might need it for specific kernel symbol lookups if `kptr_restrict` isn't fully opened up on the host. Seen it happen.

Your SELinux check is key. On AL2, it's not just `getenforce`. If it's enforcing, you need the exact policy type: `sestatus`. Sometimes it's `container_t` causing the block even with a privileged pod.


Build once, deploy everywhere


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

Adding `BPF` is indeed the critical move, but I disagree on your suggestion to drop `SYS_MODULE`. On kernel 5.10, with certain Amazon Linux 2 hardening profiles, the agent's need to resolve kernel symbols for proper probe attachment often necessitates `SYS_MODULE` as a fallback, even with privileged mode. It's not about loading modules; it's about bypassing `kptr_restrict` through an alternate syscall path.

The `getenforce` check is a good start, but it's insufficient. The policy module matters more. On EKS nodes, you're typically dealing with `container_t` or `svirt_lxc_net_t` contexts, and a privileged pod might still be blocked from creating eBPF maps if the policy hasn't been explicitly amended. `sestatus -v` or checking audit logs is the only reliable way to confirm a denial.


—BJ


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

You've got the standard config, but that's the problem - it's just the starting point. On kernel 5.10 with a hardened AL2 base, privileged and the big three capabilities often aren't enough.

The error points to the eBPF loader. You need to explicitly add the BPF capability to your list. I'm skeptical that SYS_MODULE is necessary here, despite what others are saying; it's a massive privilege escalation for what should be a data collection task. If the agent truly requires it to bypass kptr_restrict, that's a design flaw you should note in your evaluation.

Before you add more privileges, check the host's actual security state. Run `sestatus` on a failing node. If SELinux is enforcing, the audit logs will tell you the real story, not the pod logs.



   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Ah, that snippet nails it. You've got the standard config, but `privileged: true` plus those three capabilities isn't cutting it on AL2 kernel 5.10. The `BPF` capability is missing, which is almost certainly the direct cause of the "permission denied" on the eBPF load.

Quick suggestion: add `BPF` to your `capabilities.add` list and redeploy the DaemonSet. That's solved this exact error for me in three different clusters on that same kernel version.

Curious, do you have a GitOps flow for these agent manifests? A PR with that change would be a perfect case for a small validation step in your pipeline.


git push and pray


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

You're spot on about adding the BPF capability being the most likely fix. That exact change got our manifests working when we hit the same wall.

But I'm with user23 on being wary of SYS_MODULE. If the agent's design leans on that to bypass host hardening, it's a pretty big red flag for a security tool. Maybe we should be asking if there's a way to adjust the host's kptr_restrict setting instead of piling on pod privileges?

Great point about the GitOps pipeline too. We added a simple conftest policy to flag missing capabilities like BPF in any DaemonSet spec. It's caught a few oversights already.


null


   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Ah, the old "fix the container, not the host" reflex. I get it, but leaning on pod privileges to bypass kernel hardening is a vendor's game. It lets them say "works on your hardened nodes" while quietly erasing the security posture you paid for.

You're right to be wary of SYS_MODULE, but I'd push back on adjusting kptr_restrict on the host. That's another concession. The real question is whether the agent's architecture is forcing you to choose between it working and your node actually being hardened. That's a pretty poor choice for a security product.

Conftest policies are a smart move, though. They'll catch the missing BPF, but they won't flag when you're being nudged into a dangerously permissive config just to make a third-party tool run.


Beware of free tiers


   
ReplyQuote
(@hugob)
Estimable Member
Joined: 2 months ago
Posts: 196
 

You're right on the money about the default config being a starting point that fails in the real world. I've danced that dance on three clusters now.

That skepticism on SYS_MODULE is healthy, but I've had to add it back in more times than I'd like. It feels dirty, but sometimes it's the only path forward when you're up against a hardened AMI and the vendor hasn't optimized for it. It's less about loading modules and more about the kernel symbol lookup path, which is its own kind of frustrating.

And you nailed the real fix - the host logs. I spent a whole afternoon chasing pod errors before I finally ssh'd in and checked the audit log. The denial message there gave me the exact SELinux context mismatch. The pod logs just said "permission denied," which is about as useful as a screen door on a submarine. Always check the host first.


hugo


   
ReplyQuote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

Exactly. That "choice" is the whole frustration. It feels like buying a seatbelt that only works if you first remove the airbags.

I've documented three agents for our beta program now, and two of them follow this pattern. They tout "zero-trust" but their manifests quietly ask for `privileged: true` and `SYS_MODULE` once you hit a hardened kernel. The vendor response is always to escalate privileges, never to adapt their probe method.

Your conftest point is smart, but yeah, it only catches missing pieces, not creeping permissiveness. Maybe we need a *maximum* capability rule, not just a required one.


edge cases matter


   
ReplyQuote