Skip to content
Notifications
Clear all

Guide: Setting up exclusions for our dev team's build servers without creating blind spots.

1 Posts
1 Users
0 Reactions
21 Views
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
Topic starter   [#27912]

Alright, let's wade into the inevitable morass of endpoint exclusions. The dev team is screaming because Elastic is "killing their build times" and "eating all the CPU." The immediate solution, shouted from the mountaintop, is to "just exclude the build directories and processes." And if we do that blindly, we might as well just uninstall the agent and paint a target on the server. The goal here isn't to appease the devs by creating a security black hole; it's to surgically neuter the performance impact while keeping *some* semblance of visibility.

First, let's establish what we're actually trying to exclude. It's rarely just a path. It's a combination of:
* **Noisy file events:** Hundreds of temporary `.o`, `.class`, `.tmp` files being created and deleted per second.
* **Process lineage:** The `make`, `mvn`, `gradle`, `go build`, or `npm` processes and their myriad child processes (compilers, linkers, packers).
* **Memory scans:** The agent deciding to peek inside the JVM heap of a running build process because it looks "suspicious."

A naive exclusion list that just dumps `/home/dev/builds/**` into the `files.paths` exclusions is a great way for malware to drop a payload there and call it a day. So we have to be more cunning.

Here's a snippet from a Terraform module we use to configure the Elastic agent policy. Notice we're not just throwing paths around. We're combining process exclusions with path exclusions, and we're *attempting* to keep some logging on for parent processes.

```hcl
resource "elasticstack_fleet_agent_policy" "build_agents" {
name = "build-server-policy"

config = jsonencode({
"inputs" : [
{
"type" : "endpoint",
"policy_template" : "endpoint",
"enabled" : true,
"streams" : [],
"config" : {
"policy" : {
"windows" : {},
"mac" : {},
"linux" : {
"events" : {
"file" : true,
"process" : true,
"network" : true
},
"malware" : {
"mode" : "enabled"
}
}
},
"preset" : "strict",
"exceptions" : {
"process" : [
{
"name" : "mvn",
"reason" : "Build tool - high child process churn"
},
{
"name" : "go",
"reason" : "Compiler - high file/process churn"
}
],
"file" : [
{
"path" : "/tmp/bazel-output/**",
"reason" : "Bazel ephemeral output"
},
{
"path" : "/home/jenkins/workspace/**/*.tmp",
"reason" : "Temporary build artifacts"
}
]
}
}
}
]
})
}
```

But here's the sardonic truth: this is a starting point, not a solution. You will deploy this and the devs will still complain. Why? Because the real CPU drain often comes from the real-time *analysis*, not just the logging. You've told it not to log the file events, but the kernel module or ETW provider might still be collecting them, and the agent might still be trying to correlate them. You might need to dive into the advanced policy JSON to throttle event rates or exclude specific process *command line* patterns (e.g., `go build -o /tmp/...`).

The process is:
1. Enable the policy on a single build server.
2. Let the agent logs (`/opt/Elastic/Agent/data/elastic-agent-*/logs/`) fill up with warnings about throttling or dropped events.
3. Use the Kibana detection engine to see what, if anything, is still being generated from the build paths. Is it just "info" level? Fine. Is it still generating "malware" alerts? Not fine.
4. Iterate, while constantly reminding the dev team that every exclusion is a calculated risk you're documenting for the security audit they'll ignore until there's a breach.

The blind spot isn't where you stop logging; it's where you stop *understanding* what you've stopped logging. If you can't articulate exactly what signals you're missing and have a compensating control (like network egress monitoring on the build VLAN, or immutability of the base image), you've just traded a performance problem for a security incident.

-- cynical ops


Your k8s cluster is 40% idle.


   
Quote