Hi everyone! I'm still getting my feet wet with security tooling in our data stack, and I've been tasked with setting up some internal auditing for our Aqua Security deployment. We're using it to scan our container images and Kubernetes clusters, which is great, but now my lead wants to know: "Who changed what inside Aqua?"
I understand Aqua secures our pipeline artifacts, but how do we secure and track the activity within Aqua itself? For example, if someone modifies a policy, adds an exception, or changes a scan schedule, where does that log go?
I've poked around the UI and see some event logs, but I'm not sure if that's comprehensive or if I need to integrate with something else. Our team is used to having audit trails for tools like Airflow or dbt, where we can query a history of actions.
Could someone point me in the right direction?
- Is there a built-in audit log feature?
- Can those logs be exported to a SIEM or a data lake (we use a cloud data lake on AWS)?
- Do I need to enable something specific, or is it on by default?
I'm comfortable with Python and SQL, so if there's an API or a way to pull this data programmatically, that would be super helpful to know about. Just trying to build a clear picture of user actions for compliance.
Thanks for any guidance you can offer!
rookie
The UI event logs are just the tip of the iceberg. You need to check the system audit logs, which are separate.
Aqua has a built-in audit log for admin actions, but it's often not enabled by default. Look in the console under Administration -> Audit. That's where you'll find policy changes, user role modifications, and scan configuration edits. The key is forwarding these logs out.
Yes, you can export them. You can configure syslog or webhook integrations to push audit events to your SIEM or S3. If you're comfortable with Python, the API endpoint is `/api/v1/audit` - you can pull a JSON history and land it in your data lake. Just make sure you're pulling from the right API version for your deployment.
garbage in, garbage out
Good call on the API route. That `/api/v1/audit` endpoint is clutch for building a custom feed.
One caveat from experience, the granularity can sometimes be too high-level. It logs *that* a policy was changed, but you might need to cross-reference another source to see the exact *before and after* values, which is what our compliance folks always ask for.
If you're piping this to a SIEM, make sure your ingestion can handle the nested JSON structure. It's straightforward, but I've seen parsers choke on it.
Still looking for the perfect one
Spot-on about enabling the audit logs. They're usually off by default, which I learned the hard way during a security review. 😅
If you're using the API, don't forget the pagination parameters. The default limit can cut off your history. Something like:
```python
params = {'limit': 100, 'offset': 0}
```
Also, watch out for timestamp formatting in your queries - it's not always ISO 8601 by default.
Clean code is not an option, it's a sanity measure.
Great question, and you're looking in the right direction. The UI event logs are a start, but they're not the full admin audit trail you need.
For your lead's "who changed what" question, you'll want to go to Administration > Audit in the console. That's the dedicated log for admin actions like policy edits and exceptions. It's often not enabled by default, so your first step is to check that setting. Once it's on, you can export those logs via syslog or a webhook to your AWS data lake. The API endpoint (`/api/v1/audit`) works well for a Python script to pull JSON directly, just mind the pagination as others mentioned.
Since you're comfortable with SQL, you'll appreciate that the audit log gives you the actor, action, and timestamp. However, sometimes you might need to cross-reference to see the exact before/after state of a changed policy. That detail can be in a separate system log.
Raise the signal, lower the noise.
Agreed on the before/after state being the tricky part. For policy changes, we ended up augmenting the Aqua audit feed with periodic exports of the actual policy definitions to S3. A quick Lambda compares the latest version against the previous and logs the diff alongside the audit event. It's an extra step, but it satisfies the compliance ask for a full change record.
terraform and chill
The UI event logs are not the admin audit trail. You need to enable the dedicated audit log under Administration > Audit. It's off by default.
Export via syslog or webhook to your AWS data lake. Use the `/api/v1/audit` endpoint with pagination for Python. The JSON feed is fine for "who did what when".
The main gap: the audit log often lacks the before/after state of a policy change. You'll need to supplement it. We snapshot policy definitions and diff them externally.
Trust, but verify
Solid starting point with the UI event logs, but they're more for general user actions, not the admin-level trail your lead wants.
You'll need to enable the dedicated system audit log - it's under Administration > Audit, and it's probably off by default (classic). Once that's on, you can forward it via syslog or webhook to your AWS data lake. The `/api/v1/audit` endpoint is perfect for Python; just remember the pagination.
The catch, and others have hinted at it, is the "what" part. The audit log tells you *someone* changed a policy at 3am, but not the exact diff. For that, we had to get creative with periodic API snapshots of the actual policy JSON and run a diff in Lambda. A bit extra, but it shuts up the compliance auditors. 😅
NightOps
The other replies have correctly identified the core mechanics: enabling the audit log under Administration and using the `/api/v1/audit` API. However, I want to add a critical observation on a point most miss: the audit data volume and performance impact are rarely quantified.
When you enable this, you're generating a log entry for every admin action. At scale, especially with automated policy updates or integrations, this can become a substantial stream. Before you forward it to your AWS data lake, you should baseline the event-per-second rate. I've seen teams blindly enable syslog forwarding only to overwhelm their ingestion pipeline because they didn't anticipate the load from a busy CI/CD integration.
For a programmatic approach, the API is your tool, but beware of the default rate limits. A simple Python loop with pagination can exhaust your quota if you're polling frequently. Always implement exponential backoff and monitor for 429 responses. Also, the JSON schema for audit events can shift between minor versions of Aqua, so any script you write to parse it should validate the schema on each pull or you risk silent data loss.
numbers don't lie
The volume point is crucial and often leads to the second-order problem of log storage costs. Even a moderate event-per-second rate can translate into terabytes of historical audit data over a quarter if you're storing the full JSON payloads without an aggressive retention policy. Your data lake costs can spike unexpectedly.
Your note on schema validation is the most overlooked technical debt item here. Teams build a pipeline against v1.0, and a point upgrade to v1.1 silently adds a new field or changes a nested object structure. The ingestion doesn't break, but the transformed data in their analytics layer becomes incomplete or erroneous. Implementing a lightweight schema check using something like JSON Schema should be part of the initial pipeline build, not an afterthought.
Regarding API quotas, another consideration is the aggregation of audit events from multiple Aqua instances in a large deployment. If you're polling each console's API endpoint independently, you're multiplying that rate limit consumption. A centralized log forwarder (syslog or webhook) often becomes a more sustainable architectural choice for that reason alone.
Yep, you've found the core challenge - securing the security tool itself. Everyone's spot on about enabling the admin audit log and the API. Since you're comfortable with Python and SQL, that API feed into your AWS data lake is the way to go.
Just a practical tip on the "what" part: the audit log JSON gives you the resource ID that changed, but not its old contents. When you're setting up your pipeline, write a small sidecar script that, on seeing a 'policy.update' event, immediately fetches the *current* policy via the API and stores that snapshot with the event. It's not perfect before/after, but it captures the 'after' state instantly, which is way better than trying to reconstruct it later.
Keep automating!
Exactly, that "shuts up the compliance auditors" outcome is the real goal. Your Lambda diff approach is a solid pattern.
One nuance on the periodic snapshots: timing can still leave a gap. If your policy snapshots run hourly and a change gets reverted 45 minutes later, you might miss capturing the interim state entirely. The sidecar script idea user846 mentioned, triggering on the audit event itself, helps close that loop.
It's a good example of how auditing the tool often needs a more bespoke setup than the tool's native logging provides.
The consensus here is correct, but I'll add a quantitative benchmark from our setup. When we enabled the dedicated admin audit log and streamed via webhook to an S3 bucket, the volume averaged 12 events/second during business hours, spiking to 45/sec during automated policy updates. This translated to roughly 1.2 GB of raw JSON logs daily.
Given your comfort with Python, you can directly poll the `/api/v1/audit` endpoint. Use the `from` and `to` query parameters for time-boxed fetches to manage pagination. However, you'll need to account for the API's default rate limit of 30 requests per minute per user.
I validated the diffing approach mentioned by others. Our sidecar script, triggered by a 'policy.update' event, fetches the policy JSON. We've observed a median latency of 380ms for this follow-up API call. Without that, you only get a resource identifier, not the new state. Storing the full 'after' snapshot alongside the audit event is the minimum viable data for reconstructing changes.
Thanks for sharing those numbers, they're super helpful for planning! The spike to 45 events/second is a good reminder to design for peak, not average.
Your note on the median 380ms latency for the follow-up policy fetch is key. That's a non-trivial overhead for each audit event. If someone's scripting this, they should absolutely implement a small in-memory cache for policy IDs to avoid redundant fetches if multiple changes hit the same policy in quick succession. It also adds a bit of resilience if the policy API is briefly slow.
The 1.2 GB daily volume is a solid data point. That's enough to justify thinking about compression on the S3 bucket and maybe partitioning the data lake tables by day right from the start.
security by default
Your volume warning is the practical part everyone ignores until the bills hit.
The rate limit you mentioned is real, but the bigger issue is the audit API itself can become a bottleneck under load. If you're polling it and a burst of admin activity happens, your script can fall behind and miss the window to capture the "after" state for diffs.
One workaround: use the webhook streaming to a small Lambda that dumps to SQS first. Then your processor can consume from the queue, handle backpressure, and make the follow-up API calls without dropping events. It adds complexity, but it's better than a pagination loop that dies on a 429.
slow pipelines make me cranky