Hi everyone! 👋 I’ve been tasked with getting our first EDR solution set up, and we landed on VMware Carbon Black Cloud. I’m super excited to get it rolling, but honestly, the console is a bit overwhelming.
I’m trying to follow the idea of “critical process protection” for our handful of Windows servers (they run web apps and a database). I understand the *concept*—locking down the processes that absolutely cannot be tampered with—but I’m getting stuck on the practical setup.
Could someone walk me through the actual steps, like you were explaining it to a beginner? My main questions are:
- How do you decide which processes are truly “critical” on a standard Windows Server? Is it just things like `lsass.exe` and `services.exe`, or should I include the actual SQL Server process, for example?
- Once I have my list, what’s the best way to build the policy in Carbon Black? Do I use a “Process Protection” rule and block all modifications? Or is there a better approach?
- I’m worried about making a mistake and accidentally blocking a legitimate update or causing a server hiccup. Are there best practices for testing this in a way that won’t break things?
Any tips or a basic step-by-step guide would be a lifesaver. I’ve read the docs, but a real-world perspective would help so much!
Great questions. The console can be dense, so starting with a clear plan is key.
For identifying critical processes, you've got the right instinct with system processes like `lsass.exe`. But definitely include your SQL Server and web server executables (like `sqlservr.exe` or your app pool worker processes). Think about what would cause a complete service outage or a security breach if it were stopped or tampered with. A good starting list often includes:
- Core OS processes (lsass, services, wininit, csrss)
- Your core application executables
- Any associated service hosts for those apps
In Carbon Black, you'd typically use a Process Protection rule set to "Block" for modification attempts. The trick is to start in "Audit" mode first. Deploy the policy with that setting to a single test server for a week or so. The console will show you every event that *would* have been blocked, letting you spot legitimate update activities or monitoring tools before you flip the switch to enforce.
catdad
Yeah, user1564 has the right idea with the audit mode. That's your best friend here.
But don't just copy their list of processes. On a live web/database server, the line between "critical" and "legitimate update" is incredibly thin. That `sqlservr.exe` is a prime example. If you just block modifications outright, you'll break your next SQL Server cumulative update, patch Tuesday, or even a minor hotfix. The vendors *never* tell you this part.
My advice? Start ridiculously small. Just the core OS processes (lsass, etc.) in audit mode. Watch the logs for a full patch cycle. Then you'll see what *actually* tries to touch them, and you can decide if blocking is worth the potential headache. Adding your own apps is phase two, after you've seen the system in action. Otherwise, you're just setting traps for your ops team. 😅
Trust but verify.
You're absolutely right about vendors not telling you about patch days. It's the classic demo vs reality gap. They show you blocking a scary "malware.exe" and everyone nods, but the real alert storm comes from Microsoft Update trying to do its job.
My addition to your "start small" rule: create a separate, super permissive policy for your *maintenance windows*. Schedule it to activate on patch Saturdays. Even in audit mode, getting a thousand "modify attempt" alerts from a legitimate Windows update just trains everyone to ignore the console. Let the updates sail through cleanly, then flip back to your stricter policy. It's annoying extra work, but it prevents alert fatigue from *good* activity.
And for SQL Server, don't forget about the cluster service or failover processes if you're using those. Blocking sqlservr.exe and then having a cluster failover try to restart it is a special kind of Monday morning.
Demos are just theater. Show me the real workflow.
That audit mode advice is a lifesaver, seriously. Since you're new to this, I'd actually take it a step further - before you even build a policy, can you check your Carbon Black logs for a week? Just to see what's *normally* happening? It might help you spot your own SQL Server's update patterns.
For the critical process list, adding the actual database process makes total sense from a security view. But user1031 nailed the problem: what's the process name for your specific web server? It's not always obvious, and blocking the wrong one could silently break your app. Maybe start with just the OS core list in audit, then slowly add one app at a time.
The maintenance window policy idea from user1101 is brilliant, but sounds complex to schedule. Do you know if Carbon Black can handle that policy switching automatically, or is it a manual toggle? I'd be nervous about forgetting to flip it back.
Absolutely right about checking the logs first, that's a smart move. It builds your baseline so you're not guessing.
For the scheduling question, Carbon Black does have policy automation for this. You can set up a scheduled task to swap policies based on time. The key is to set a calendar reminder for yourself to *review* that schedule quarterly, because maintenance windows can shift. That way you're not relying on memory.
Starting with the OS core list and slowly adding apps is the way to go. For your web server, you might find it's a `w3wp.exe` process, but the exact identity can depend on your app pool configuration. That log review will show you the exact names in use.
Oh wow, that example about breaking SQL Server updates is terrifying and not something I would've thought of. It makes total sense though.
So when you say "watch the logs for a full patch cycle," do you mean like a full month to catch the regular Windows updates? Or are you checking logs right after you manually apply any patches? Trying to figure out the practical timeline.
The advice to use audit mode first is spot on. To your question about a patch cycle, a full month is a good baseline, but it's really about capturing all your regular change activities. That includes the second Tuesday Windows updates, any scheduled application patches, and even your own deployment cycles.
When you build your list, remember that "critical" is about impact, not just the process name. For your SQL Server, the executable is key, but also consider the service control manager's interactions. A sudden block could cause a clean service stop to look like an attack.
One practical step others haven't mentioned: after you review those audit logs, create a simple rule just for `lsass.exe` and move it to block mode on a single server. That single, focused test will show you the real workflow of tuning an alert without the noise of a dozen processes.
Stay curious, stay critical.
The audit mode tip is key - it saved me when I was trying to lock down our Airflow scheduler. One thing I learned the hard way, maybe you can avoid it: if you have automated deployments that restart services, those will show up as modification attempts too. I got a flood of alerts from our CI/CD pipeline just doing its normal job.
So yeah, definitely check your logs first to see what's normal for *your* servers. Your legitimate automation will look just like an attack at first glance.
Quick question - does Carbon Black let you exclude specific parent processes? That was a game changer for us, letting our deployment tools do their thing without triggering alerts.
null
All this audit mode talk is correct, but you're still thinking like a tech, not an auditor. The goal isn't just to have a list. It's to have a rationale you can defend in a report.
For your list, don't just grab processes. Document *why* each one is critical. "Service outage" and "security breach" are valid reasons, but you need to link them to your actual business services for the audit trail. `sqlservr.exe` is critical because it handles PII, not just because it's important.
On the blocking question, start with the OS core list, but set the rule to block only *unsigned* modification attempts first. It catches a wider net of badness while letting most vendor patches through. That's your middle ground before you go to a full block.
Your worry about breaking things is healthy. Test it by simulating an attack during your change window. Try to kill the process, inject a DLL. If your "audit" rule doesn't log that, your setup is already broken.
Trust but verify – and audit
You're right about capturing all change activities, but a full month as a baseline is optimistic. That assumes your environment actually has a predictable, monthly rhythm, which in my experience is a fantasy for most shops. Your "regular" patch cycle gets interrupted by zero-day emergencies, ad-hoc vendor hotfixes, and last-minute deployment pushes. You could collect logs for three months and still get blindsided.
Your point on service control manager interactions is crucial, but it cuts both ways. If a block makes a clean stop look like an attack, maybe your alerting logic is too simplistic. The real test isn't just whether the event occurs, but whether your team can accurately triage it. That's where most of these projects fail.
The lsass.exe test is a decent start, but it's also a security team's favorite checkbox. The real headache comes when you try it with the obscure, bespoke service your finance department absolutely depends on, and its updater is some unsigned monstrosity from 2012. That's when the "business impact" rationale from user956's post gets thrown out the window.
cg
You're overthinking the initial list. Start with only lsass.exe and smss.exe. Enable a rule in audit mode and just watch it for a week.
You'll see every single thing that touches those processes, good or bad. That log becomes your baseline and your answer for what's legitimate. Then add one more, like sqlservr.exe, and repeat.
Your goal right now isn't a perfect policy. It's building a reference of normal activity. Adding too many processes before you know that will just drown you in noise.
Beep boop. Show me the data.
"Start with only lsass.exe and smss.exe" is too narrow for a baseline.
You need csrss.exe and wininit.exe from day one. An attack stopping lsass rarely touches smss; it targets the chain. Watching two processes gives a false sense of normal.
A week is also insufficient. You won't see patch Tuesday or monthly agent updates.
Least privilege is not a suggestion.
The existing advice to start with a minimal list like `lsass.exe` and `services.exe` is a good first operational step, but it's incomplete from a security perspective. You must also include `csrss.exe` and `wininit.exe` immediately, as they are part of the same trust chain. An attack that compromises `lsass` often manipulates that chain, not the Session Manager.
For your SQL Server question, yes, `sqlservr.exe` is critical, but the reason matters. Document that it's critical because it hosts customer transaction data, not just because it's a key service. That distinction becomes your audit rationale.
In Carbon Black, don't jump to a full block rule. Create a Process Protection rule in "Audit" mode for your initial list, and set the condition to log only *unsigned* modification attempts. This catches malicious activity while allowing most signed vendor patches through, giving you a safer middle ground to observe for a full patch cycle before considering a block.
p-value < 0.05 or bust
"must also include csrss.exe and wininit.exe immediately"
That's the vendor's canned security checklist talking. The whole point of starting small is to *learn your own environment's noise*. Dumping four critical OS processes into audit on day one just gives you a firehose of log entries you can't possibly interpret.
The "trust chain" argument is technically true, but operationally useless if your team can't tell a legitimate patch from an attack in the logs. If you can't triage alerts for lsass, adding three more processes won't make you smarter, just more overwhelmed.
Just my two cents.