Skip to content
Notifications
Clear all

Did anyone get Aqua Security to work well with GitLab CI?

2 Posts
2 Users
0 Reactions
23 Views
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
Topic starter   [#18463]

I have been evaluating Aqua Security's vulnerability scanning and compliance tools for integration into our GitLab CI/CD pipelines over the past quarter. The stated goal is to shift security left without introducing significant friction or unpredictable cost overhead into our deployment cycles. While the Aqua documentation provides a foundational overview, I have encountered several operational complexities that merit a detailed community discussion, particularly concerning performance and the associated cloud expenditure.

My primary configuration involves using the Aqua Trivy scanner as a GitLab CI job, invoked via Docker-in-Docker. The pipeline is designed to scan both our application container images and the underlying infrastructure-as-code templates. A simplified version of the job definition is as follows:

```yaml
container_scan:
stage: test
image: docker:stable
variables:
DOCKER_DRIVER: overlay2
TRIVY_VERSION: 0.45.0
services:
- docker:stable-dind
script:
- docker run --rm -v /var/run/docker.sock:/var/run/docker.sock
-v $CI_PROJECT_DIR:/src aquasec/trivy:${TRIVY_VERSION}
image --severity CRITICAL,HIGH --exit-code 1 --format table
--input /src/myapp.tar
allow_failure: false
```

The functional challenges are multifaceted:
* **Scan Duration:** Scan times for our mid-sized Node.js and Java applications range from 4 to 7 minutes per image. This adds a non-trivial delay to our pipeline execution, which compounds across multiple parallel pipelines. I have observed that the database updates for vulnerability definitions can cause significant variance in job runtime.
* **Resource Consumption:** The DinD (Docker-in-Docker) approach, while functional, consumes considerable compute resources on our GitLab runners. Our AWS EC2 (c5.2xlarge) runner instances show a 40-50% increase in average CPU utilization during scans, directly impacting our ability to run concurrent jobs and forcing consideration of auto-scaling groups, which introduces cost volatility.
* **Cost Attribution:** The compute time for these security scans is currently absorbed into a general "CI/CD" cost center. I am attempting to isolate the exact cost of the Aqua scanning step—including runner compute, ECR pulls for the scanner image, and network egress—to justify its value and optimize its scheduling. Preliminary estimates suggest it adds 15-20% to our monthly CI compute bill.

My specific inquiries for the community are thus quantitative and architectural:
* Has anyone successfully implemented Aqua with GitLab CI in a way that minimizes pipeline latency? What specific runner configurations (instance type, caching strategies, or alternative scanning methods like the Trivy GitLab native integration) yielded measurable improvements in seconds-per-scan?
* How are you allocating the cloud costs incurred by these security scans? Are you using AWS Cost Allocation Tags, Azure Resource Groups, or a third-party FinOps tool to attribute the expense back to the security or platform engineering teams?
* Were there any non-obvious configuration parameters for Trivy—such as the `--skip-db-update` flag with a managed external database, or custom policies—that substantially reduced resource consumption without compromising scan integrity?

I am particularly interested in concrete metrics: percentage increase in pipeline duration, average cost per scan, and the breakdown of infrastructure costs pre- and post-integration. Anecdotal evidence of "it works" is less useful than specific data points on performance degradation and cost per finding.

Show me the bill.


CostCutter


   
Quote
(@dannyz)
Estimable Member
Joined: 3 months ago
Posts: 171
 

Oh, I'm actually trying to set something similar up right now. Your job definition looks way more advanced than what I've managed so far. I'm still on the basic setup.

Could you maybe share more about the "operational complexities" you hit? I'm a bit nervous about the performance part myself. Does it really slow down your pipeline a lot?



   
ReplyQuote