Skip to content
Notifications
Clear all

Guide: Setting up critical process protection for servers, step by step.

42 Posts
40 Users
0 Reactions
19 Views
(@amymk)
Estimable Member
Joined: 2 months ago
Posts: 115
 

That's exactly the question I had when we set ours up last month. The advice to skip the list and just start logging on one test server saved us.

We did it on a staging database server first. We created the broad audit rule and let it run. It was shocking how many things touched the SQL process that we never thought about - the backup software, the monitoring agent, even a logging service. If we'd just blocked everything on our 'critical list' right away, our nightly backups would have failed.

So my tip? Start with that audit rule on a single, non-essential server. Let it run for at least a full week to catch all the regular tasks. The noise is annoying, but it shows you what's actually normal so you don't break it later.



   
ReplyQuote
(@carlam)
Reputable Member
Joined: 3 months ago
Posts: 234
 

Totally agree about treating license start as day one. That said, I've found some vendors make you pay extra to retain historical demo data, which feels like a trap for teams who get attached to their early work.

But there's a practical upside to starting fresh: your alert thresholds and baselines are calibrated against real, stable operations from the start. No weird spikes from your test period screwing up the averages.


Benchmarking my way to better decisions


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

The advice to start with a single test server in audit-only mode is the correct path. You've hit on the exact risk everyone faces: that initial "critical list" you build in a vacuum will almost certainly miss legitimate processes, like your backup software or monitoring agent touching the SQL Server executable.

Your specific question about the SQL process is key. In Carbon Black, you would use a Process Protection rule for something like sqlservr.exe, but you shouldn't start by blocking all file modifications. That's how you break patching. Focus the rule's operations on runtime threats like process injection, hollowing, or tampering. That secures the running service while still allowing your approved tools to update the binary on disk.

Run that broad audit rule for a full business cycle - including patch windows and backups - on your test box first. The log noise is frustrating, but it's the only reliable way to see the actual parents and service accounts that need to be in your exclusions. Only then should you consider moving a policy to block mode.


Review first, buy later.


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

That's the right approach, but be careful with the `Get-Process` baseline over a week. On its own, it's a flat snapshot. You'll see parents, but you'll miss the *context* of *when* those parents spawn the child process, which is where the real pattern hides.

Better to pair it with a scheduled task that logs the parent *and the command line* every 30 minutes. You'll catch that the monitoring agent restarts sqlservr.exe with a specific config flag only on patch day. The parent process alone won't tell you that.


- elle


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

Good point about the context. So you're saying a scheduled task with command line logging gives you the "why" behind the process launch.

But doesn't logging command lines every 30 minutes still miss the moments in between? What if a short-lived malicious process spawns right after one of your snapshots?



   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Great questions. Everyone's nailed the main point: ditch the list and start in audit mode.

Your worry about blocking updates is exactly why. On our first run, we nearly broke a core accounting process because we didn't know the inventory tool needed to spawn it. Let the audit rule on a non-prod server collect data for a full week, patches and all. You'll see all the legitimate parents for things like `sqlservr.exe`.

Then, for your protection rule, target runtime actions like injection, not file modifications. That locks down the live process while letting your patch tools do their job.


Trust the trial period.


   
ReplyQuote
(@fred99)
Estimable Member
Joined: 3 months ago
Posts: 95
 

The audit mode advice is the only thing that saved me when I was in your exact spot. The real learning wasn't the list, it was seeing all the unexpected legitimate activity, like our AV scanner and performance monitor.

One thing I'd add: when you look at your audit logs, pay close attention to the "operation" field. You'll quickly see which actions are almost always malicious (like injection) versus normal file writes from patching. That's how you know what to actually block later.

How comfortable are you with parsing the Carbon Black event logs after a few days of audit data?



   
ReplyQuote
(@helenb)
Estimable Member
Joined: 3 months ago
Posts: 128
 

The manual toggle risk is real. I've seen similar rules in other compliance systems where forgetting to revert after a maintenance window causes audit failures the next day.

Does Carbon Black have an API for that? You could theoretically link the policy switch to your change management ticket system, so it flips back when the ticket closes.

Your point about slowly adding one app at a time is how we built our financial software controls. We started with the core OS, then added the main accounting process, then the reporting service. Each one ran in audit for a full month-end cycle before we enabled blocking. It exposed a legacy integration we'd forgotten about.



   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

The point about `csrss.exe` and `wininit.exe` is critical, but I'd add a deployment caveat. In large environments, especially those with legacy vendor software, you sometimes find poorly coded applications with hooks into those core Windows processes. Including them in an initial audit can flag a huge volume of noise.

The better approach is to scope the initial audit rule to servers in a specific, modern workload tier first. Apply the full list, including those core processes, to a pilot group of, say, your Kubernetes nodes or front-end web servers where the process interaction model is cleaner. Then expand to more complex application servers in a second wave. This isolates the signal from the legacy tool noise.


—Alex


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

Lots of good advice in the thread, especially the heavy emphasis on audit mode. But honestly, starting with a "list" of processes is your first trap.

You ask how to decide what's critical. You don't. Your server's normal operation does. If you sit down and brainstorm it, you'll miss the weird vendor monitoring agent or the backup software's obscure child process. Build your list from the audit data, not the other way around.

That said, your SQL Server process is absolutely a candidate, but not by blindly blocking "all modifications." You'll torpedo every patch Tuesday. Focus the protection rule on things like process injection, memory tampering, or hollowing - the runtime attacks that are almost never legit. Let your patch manager write to the sqlservr.exe file on disk, but lock down the nasty stuff that happens after it's running.

Test it? The best practice is to not enable any blocking at all for at least a full business cycle. Let it run in audit on a non-critical but identical server. You'll be shocked at the noise.


Trust but verify.


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

You've hit the nail on the head about thinking like an auditor. I'd push it one step further: the "why" for each process should map directly to a risk register item or a compliance control (like PCI DSS 6.3.2 or whatever framework you're under). That turns your Carbon Black policy from a technical config into an audit artifact.

The unsigned-block first approach is smart, but I've seen it trip up old-but-still-supported vendor software that uses self-signed or expired certs on their installers. It's a good middle ground, but your audit logs better include the hash and signer details so you can make the exception *before* you switch to block mode.

Simulating attacks during your change window is brilliant. It validates your detection logic when you know the system is already in a known-good state.


Sleep is for the weak


   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Exactly. That scheduled task log becomes your baseline's timeline, which is gold for alerting later. You can pipe it into Prometheus to track parent process spawn rates. A sudden spike in `sqlservr.exe` launches from `YourMonitoringAgent.exe` outside the patch window is a perfect alert condition.

But you'll want to aggregate those logs aggressively, or you'll drown in noise. I'd sum them up by hour and parent process name before feeding them to the alert rule.


Sleep is for the weak


   
ReplyQuote
Page 3 / 3