Skip to content
Notifications
Clear all

How to stop Lacework from scanning our development/transient environments?

15 Posts
15 Users
0 Reactions
11 Views
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
Topic starter   [#26100]

Our organization has been running Lacework for container and cloud security posture management across our entire AWS footprint for approximately 14 months. While the value proposition for production and staging environments is clear, the operational and financial overhead of its continuous scanning in our development and transient environments has become a significant point of contention. These environments, which include short-lived feature branch deployments, developer sandbox accounts, and integration testing clusters, are characterized by high volatility and intentionally permissive configurations. The volume of findings they generate is overwhelming our security team's capacity for triage and, based on our analysis, is directly correlating with a measurable increase in our monthly Lacework consumption costs.

The core issue is that Lacework, by default, applies its full suite of agent-based and agentless assessments to any cloud account or Kubernetes cluster once it is integrated. Our objective is to implement a precise, policy-driven exclusion mechanism that prevents Lacework from scanning specific resources or entire environments without disabling the platform for our critical workloads. We have explored the documentation but require a concrete, battle-tested strategy.

Our current multi-account AWS structure is as follows:
- `aws-123456-prod` (Production)
- `aws-789012-staging` (Staging)
- `aws-345678-dev` (Permanent Dev Resources)
- `aws-901234-feature-*` (Ephemeral Feature Branch Accounts, auto-provisioned)
- `aws-567890-sandbox-*` (Individual Developer Sandboxes)

We need to exclude all resources within the `feature-*` and `sandbox-*` account patterns, as well as specific tagged resources (e.g., `Env: transient`, `Purpose: testing`) within the permanent dev account. The desired outcome is zero data ingestion, compliance evaluations, or vulnerability scans originating from these excluded scopes.

We have attempted the following with mixed results:

* **Lacework Resource Groups:** Created groups using AWS account ID exclusions and tag-based rules. While this organizes the console view, our logs indicate scanning activity continues unabated on the excluded resources, suggesting these groups are for filtering *displayed* data, not for halting *collection*.
* **Agent Configuration (`config.json`):** For hosts where we run the container agent, we tried scoping the `acceptedHosts` field. This is untenable for ephemeral containers and does not address agentless cloud resource assessments.
* **AWS SCPs / IAM Policies:** We considered denying Lacework's IAM role specific actions in target accounts, but this risks creating configuration drift and alert noise if the role is partially crippled.

The most promising avenue appears to be the **Lacework Policies** module, specifically creating suppression rules. However, the granularity required—suppressing all alerts for entire AWS accounts—seems to conflict with the event-based, finding-specific design of the policy engine.

Could the community provide a detailed, operational blueprint for achieving this? Ideal responses would include:

1. The definitive Lacework component (e.g., Resource Groups, Policies, API endpoints) that acts as a true *collection* filter, not just a *view* filter.
2. Specific configuration examples for achieving account-level and tag-based exclusions. For instance, a valid Terraform snippet for the `lacework_integration_aws_ct` resource or a direct API call to the `/api/v2/CloudAccounts` endpoint.
3. Any mandatory ordering of operations (e.g., must create suppression before integration, must use tags applied at resource creation).
4. Empirical data on resultant cost reduction or reduction in finding volume from those who have implemented similar exclusions.

Our stack is primarily Terraform-managed, so solutions compatible with Infrastructure as Code are strongly preferred. We are operating on the latest Lacework agent version and use both AWS CloudTrail integration and the container vulnerability scanner.



   
Quote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

You've perfectly identified the core architectural limitation. Lacework's default 'all or nothing' posture for integrated accounts is a common pain point. While their documentation suggests using Resource Groups or Agent Access Keys for scoping, in practice these are blunt instruments.

Our team solved this by shifting the control plane from Lacework to our own IaC. We now tag every AWS resource and EKS cluster with an environment-specific tag (e.g., `env=transient`). In Terraform, the Lacework provider integration includes conditional logic based on that tag. For transient environments, we simply don't deploy the Lacework agent Helm chart or the CloudTrail integration module at all. This creates a clean, automated exclusion.

The caveat is this only works if you have full IaC control. For existing ad-hoc accounts, you'll need to use Lacework's API to programmatically disable the integrations, which is more cumbersome. The cost correlation you're seeing is real; their pricing model does not distinguish between a critical production alert and noise from a sandbox.


—Alex


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Oh, the classic "scan everything, bill for everything" model. Your cost correlation isn't a guess, it's a fact - more resources, more events, more ingested data equals a bigger invoice. Been there.

The IaC exclusion method user1436 mentioned is the right idea for control, but if you need something quicker and don't have full Terraform coverage, try this: use AWS Resource Groups and Tag Editor to bulk-tag all resources in a dev account with something like `lacework:scan=false`. Then, in Lacework, you can create a Compliance Exception policy at the *organization* level that suppresses alerts for resources with that tag. It doesn't stop the data ingestion (so cost remains), but it silences the alert noise drowning your team.

To actually cut the cost, you need to stop the data flow. For transient AWS accounts, the nuclear option is to simply *not* onboard them to Lacework in the first place. Use a separate, dedicated OUs in AWS Organizations for your ephemeral stuff and exclude that entire OU from your Lacework integration. Their sales team will hate it, your finance team will love it.


- elle


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Right, the compliance exception for alert noise is a solid short-term fix. We used it for a few months.

But that nuclear option of not onboarding certain accounts? That's the real cost-saver. We set up a separate AWS Organizational Unit for all our sandbox and short-lived project accounts. Our Lacework CloudFormation StackSet is scoped to exclude that OU entirely. No data ever leaves, so the bill stays predictable.

Just make sure your security policy officially defines what qualifies as a "transient" environment, so it doesn't become a loophole.


Automate the boring stuff.


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

Interesting. So you're saying the real cost driver is the data ingestion from scanning those volatile environments, not just the alert noise.

How do you handle the security review for something deployed in a dev sandbox that later gets promoted? Like, if it was never scanned until it hits staging, does that cause issues?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Exactly right, the ingestion is the killer. Alerts you can suppress, but the meter is always running on that raw data stream.

Your question about promotion is spot on - that's the operational gap this creates. We solved it by making the first deployment to a "pre-staging" environment the security checkpoint. That environment *is* fully onboarded to Lacework. Nothing gets promoted out of a sandbox, it gets rebuilt from source into pre-staging, which triggers the first real scan. It adds a small gate in the pipeline, but it's cleaner than trying to back-scan a configured resource.

The risk, of course, is a developer baking something truly nasty in their sandbox and only finding out later. You have to weigh that against the cost of scanning every experiment.


Pipeline is king.


   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

The problem you're hitting is a common misalignment between a tool's operational model and actual environment lifecycle. Lacework's default posture assumes all integrated resources have equal governance priority, which isn't true for dev/transient systems.

Your objective for a policy-driven exclusion mechanism is spot on, but you'll likely need to build it outside Lacework. The most precise method I've seen uses your CI/CD pipeline as the control point. For any deployment tagged with `env=transient`, the pipeline simply omits the step that injects the Lacework agent or registers the cloud account. This requires treating the integration as a deployable component, not a global setting.

Have you evaluated whether your transient environments truly need any runtime security scanning, or could you shift that left to static analysis in the merge request? That's often a more cost-effective boundary.


sub-100ms or bust


   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

You've correctly identified the lack of native, granular policy controls as the operational constraint. The phrase "policy-driven exclusion mechanism" is key, because Lacework's tooling simply isn't built for that level of dynamic selectivity within a single integrated account.

The most effective pattern I've documented involves a proxy layer: an internal configuration service that your deployment pipelines call to determine if Lacework integration should be provisioned for a given environment. This service evaluates your defined policies (tags, account type, lifespan) and returns a boolean. It adds complexity, but it decouples your lifecycle rules from Lacework's static models.

A critical caveat to the CI/CD control method is agent cleanup for ephemeral Kubernetes clusters. If your pipeline omits the agent installation, that's clean. But if you need to *remove* scanning from an existing cluster, you must have a separate, automated process to uninstall the agent DaemonSet, or you'll have orphaned resources reporting errors. This is often overlooked in transient environment design.



   
ReplyQuote
(@charlie2)
Reputable Member
Joined: 3 months ago
Posts: 345
 

Wow, that sounds like a huge drain on both time and budget. I'm new to Lacework, but this is a classic case of a good tool causing friction when its model doesn't match your workflow, isn't it?

You mentioned wanting a precise, policy-driven exclusion. Since you're already tagging things, what would you recommend as the most reliable signal? Something like an `env:transient` tag? Or would you hook it into the project creation process itself?

Also, have you talked to your Lacework account rep about this? I'm curious if they have any native features in the pipeline for this kind of scoping.



   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your analysis is correct - the financial pressure from scanning volatile environments is a direct consequence of Lacework's billing model, which is based on data ingestion volume.

You're looking for a policy-driven exclusion mechanism. The most precise method I've implemented uses a deployment pipeline check. For any environment tagged `env=transient` or `lifespan:ephemeral`, the pipeline script skips the Lacework agent installation entirely. This requires codifying the integration as a deployment step, not a platform-wide setting.

Have you considered a cost-benefit analysis to define a formal policy? We documented that environments with a lifespan under 48 hours and no access to production data are exempt from runtime scanning. This policy then drives the automation, whether through pipeline logic or OU exclusion, and satisfies internal audit requirements.



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

It sounds like you've hit the classic scaling pain where the tool's one-size-fits-all approach clashes with modern environment lifecycles. You're right about the default behavior - once an account is integrated, it's all in.

The policy-driven mechanism you're after is the real goal. From what you've described, the most reliable control plane for that might actually be your Infrastructure as Code (IaC) framework itself. If all your environments, especially the ephemeral ones, are spun up via Terraform or CloudFormation, you can bake the Lacework integration decision into those modules. For example, a module input variable like `enable_runtime_security = false` that conditionally excludes the agent setup and cloud account registration.

That way, the policy is defined at provisioning time, and it's completely automated. Have you looked at whether your transient environments are consistently provisioned through a single pipeline or template? That's usually the best hook for this.


ship it


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

Your point about the tag-based suppression not stopping the cost is exactly why we stopped doing it. We went down that path initially, and it created a false sense of control. The finance reports kept climbing because the data kept flowing, even though the SOC console was quiet.

The separate OU method you mentioned is the only real way to cap costs, but it introduces a significant logging gap. We had to implement a compensating control by forcing a hardened, scanned "golden image" or a specific, approved CloudFormation stack for any workload spun up in those excluded OUs. It's an extra step, but it was the trade-off for predictable billing.

Have you run into any internal audit pushback on having entire OUs with no runtime visibility?


Logs don't lie.


   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Your separate OU method is the right long-term play, but it introduces a new IaC problem. How do you keep those modules for excluded OUs in sync with the ones that have Lacework enabled? You can't just have two different stacks.

We solved it with a single Terraform module that uses a `lacework_enabled` bool variable. It defaults to true for the core modules, but our pipeline overrides it to false for any deployment tagged as ephemeral. This way, the security decision is codified at provisioning time, not an afterthought.


—cp


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

It's the classic hole in this whole approach. If your first real scan is in pre-staging, you've already baked the cake. Finding a major vulnerability then means halting a promotion and scrambling to fix something that should have been caught in a dev loop.

The risk isn't just "something truly nasty." It's a trivial misconfiguration that becomes a blocking issue at the 11th hour because your cost-saving measure created a security blind spot.


CRM is a necessary evil


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Exactly. This is the painful trade-off everyone is trying to engineer their way around. You either pay the tax for full visibility or you accept a blind spot.

But let's be honest - the "trivial misconfiguration" you're worried about? That should be caught by a shift-left tool in your commit pipeline, not a runtime agent scanning a transient environment. If you're relying on Lacework to find basic config errors in pre-staging, your CI process is already broken.

The real hole is assuming one tool has to do everything.


been there, migrated that


   
ReplyQuote