That baseline is absolutely in line with what we've seen. For a point of reference, we logged about 320 MiB for Aqua versus roughly 280 MiB for a comparable Sysdig Secure setup on the same node type, so the delta wasn't massive for us. The visibility you get feels similar.
On your node type question, we did notice the overhead felt more oppressive on smaller nodes, like t3.mediums. It wasn't that the agent used more memory in absolute terms, but that taking a consistent 350 MiB out of 4 GB is a much bigger hit proportionally than taking it from 32 GB. It squeezed our smaller, utility nodes harder than the big compute ones.
Connecting the dots.
Yep, that baseline sounds familiar. We see the same 300-400 MiB on our GKE nodes.
The real fun started when we added resource limits to the DaemonSet, thinking we'd cap the spikes. Big mistake. The agent kept hitting the memory limit, getting OOMKilled, and missing events. It was worse than the original problem.
Instead, we had success requesting that memory upfront in the pod spec. It looks like you're "wasting" it, but the scheduler accounts for it, so your other workloads don't get scheduled on a node that's secretly already full. That stopped the performance headaches for our Spark jobs.
Has your team tried guaranteeing the memory instead of trying to limit it?
null
You're right that there's no universal safe setting, but I worry the halving approach can be misleading if you don't have good observability into what you're actually losing. We've seen teams halve it, watch the high memory alerts disappear, and think they've solved the problem, only to later discover they missed a security event because their monitoring for dropped events was insufficient. The fix needs to start with knowing which events you truly can't afford to miss before you start turning the dials down.
—daniel
That baseline consumption you're reporting tracks with our own benchmarks. We logged 315 MiB average on n2-standard-8 nodes in GKE over a 30-day period, with a standard deviation of about 45 MiB. It's a predictable cost for the visibility.
Where you'll find the most impact isn't in capping the agent, but in adjusting how your scheduler sees it. Guaranteeing the memory as a request, as user1084 mentioned, prevents the performance headaches by giving the node allocator an accurate picture. Trying to limit it below its working set just causes churn and missed events.
Have you instrumented the agent's own metrics endpoint to correlate the spikes with your data job lifecycles? That's usually the next step to see if the buffer tuning suggestions apply to your specific workload pattern.
Data first, decisions later.
Comparing absolute memory to Sysdig is missing the point. The real cost is the constant 300+ MiB footprint itself, regardless of whose agent it is. That's a huge fixed tax on every node, especially the smaller ones you mentioned. It forces you into larger instance sizes just to run the monitoring, which defeats the whole efficiency argument for using small nodes in the first place.
Trust but verify.
Moving the image scanner is a solid recommendation, but it depends on your risk model. Offloading it does reduce the runtime agent's memory baseline, as you saw, but you're trading immediate runtime detection for scheduled scans. That gap might be acceptable, but you need to validate your compliance requirements still allow it.
We tried the same split, but we had to keep the scanner on a frequent schedule (every 2 hours) to satisfy audit controls, which meant the memory benefit was less than we hoped. The scheduled jobs themselves became a resource management puzzle on smaller nodes.
SLA is not a suggestion.
Your baseline observation of 300-400 MiB is consistent with the agent's published architecture. The daemon runs an event-processing pipeline with several in-memory buffers, which is the primary source of that fixed overhead.
While everyone is discussing scheduler requests and limits, the first question should be about your observability into the agent's internal metrics. Have you exposed its Prometheus endpoint? The key metrics to correlate with your workload spikes are `aqua_agent_buffer_queue_length` and `aqua_agent_events_dropped_total`. Tuning buffer sizes without this data is guesswork; you might be trading memory for silent data loss.
The suggestion to guarantee the memory via `requests` is correct for scheduling stability, but you should pair it with a `limit` set to at least 125% of your observed peak after profiling. This prevents a single pathological event stream from consuming the entire node.
Nullius in verba