Skip to content
Notifications
Clear all

Troubleshooting: Aqua's enforcer pods stuck in 'CrashLoopBackOff'.

4 Posts
4 Users
0 Reactions
20 Views
(@integrations_ivan)
Reputable Member
Joined: 7 months ago
Posts: 242
Topic starter   [#20144]

I've been architecting a deployment pipeline integration for a customer using Aqua's Kubernetes enforcers, and we've hit a persistent issue that seems to be a common pain point: enforcer pods entering a `CrashLoopBackOff` state shortly after deployment. The environment is a managed Kubernetes service (EKS), and we're following Aqua's Helm-based deployment for version 6.5.

The core symptom is that the `aqua-enforcer` DaemonSet pods crash repeatedly. Examining the pod logs before the crash yields a recurring error, but the root cause appears to vary. In our specific case, the logs indicated a failure to communicate with the Aqua Gateway, but I've also seen this happen due to permission issues, SELinux contexts, or missing kernel modules.

To facilitate a structured troubleshooting approach, I've compiled the key investigation steps we took. I'm posting this in hopes of both documenting a solution and crowdsourcing other potential root causes the community has encountered.

**Primary Investigation Paths:**

1. **Logs Analysis:** The first and most critical step. The immediate container logs often point to the specific failure.
```bash
kubectl logs -n aqua --previous
```
Common log errors we've seen include:
* `Failed to connect to gateway:443` (network/connectivity)
* `Error: failed to start container: executable file not found in $PATH` (image corruption)
* `cannot open shared object file: No such file or directory` (missing dependencies in the host OS)

2. **Host-Level Dependencies:** The enforcer requires specific kernel modules and configurations. On the affected nodes, verify:
* `overlay` and `nf_conntrack` modules are loaded (`lsmod | grep -E 'overlay|nf_conntrack'`).
* AppArmor/SELinux compatibility, especially on EKS with Bottlerocket or RHEL-based workers.

3. **Network Policy and RBAC:** The enforcer requires outbound access to the Aqua Gateway/Server on specific ports (typically 8443). It also needs a robust ClusterRole. We had to adjust our `NetworkPolicy` and ensure the enforcer's service account had the necessary permissions to `get`, `list`, and `watch` pods and nodes.

**Our Resolution:**
In our case, the issue was twofold. First, a global `NetworkPolicy` was default-deny and blocking egress to the Aqua Gateway service ClusterIP. Second, the EKS nodes were running a kernel version that required an additional `iptables` flag for compatibility. The final enforcer DaemonSet configuration snippet that resolved the crash was an addition to the `env` section:

```yaml
env:
- name: AQUA_USE_LEGACY_IPTABLES
value: "true"
- name: AQUA_SERVER
value: "aqua-gateway.aqua:8443"
```

I'm particularly interested in scenarios where the logs are less conclusive. Has anyone encountered a `CrashLoopBackOff` where the primary failure was related to **resource constraints** (memory, CPU) on the host, or specific **cloud provider security group** configurations that only manifest after the pod's initial health check? Furthermore, in an event-driven architecture where the enforcer is a critical control point, how are you monitoring its health beyond basic Kubernetes liveness probes?


Single source of truth is a myth.


   
Quote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

The logs are the obvious first step, but you're spot on that the error there is just the final symptom. In my experience on EKS, the gateway connectivity error is often a red herring.

You'll want to verify the enforcer can even reach the gateway endpoint from a node perspective, not just the API server. I've seen it where the gateway service type changed between chart versions and the internal cluster DNS wasn't resolvable from the pod's network namespace.

More critically, check the kernel module dependencies on your EKS worker nodes. If the required modules aren't present or can't load, the container will bail almost immediately, sometimes generating a generic error that gets logged as a gateway comms failure. The DaemonSet's init container logs are more telling for that.



   
ReplyQuote
(@dianaf)
Reputable Member
Joined: 3 months ago
Posts: 260
 

Totally agree on logs as the first step, but I've noticed the `--previous` flag can sometimes be useless if the pod never got to a "Running" state at all, which happens a lot with these enforcers. You might need to add `-c` for the specific container name too, especially with the init containers.

Can you share a snippet of the exact error from your logs? I'm curious if it's the classic "failed to create connection" to the gateway or something more cryptic about the runtime. We saw something similar last month that turned out to be a mismatch between the gateway service name in the enforcer config and the actual Kubernetes service DNS.



   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

That's a solid foundation for a structured approach. I'd suggest elevating the logs analysis step by explicitly separating the container logs from the initContainer logs right from the start. In Aqua's enforcer DaemonSet, the main container often crashes because a prerequisite step in the initContainer failed. You need to check both sequentially.

For example, a command like:
```bash
kubectl logs -n aqua -c enforcer-init
```
followed by
```bash
kubectl logs -n aqua -c aqua-enforcer
```
can reveal if the issue is in the module loading or host filesystem mounting (init) versus the runtime connection to the gateway (main). I've seen the initContainer fail due to a missing `kernel-headers` package on the underlying node, which only shows up in its logs, while the main container's logs just show a generic crash.


Extract, transform, trust


   
ReplyQuote