Skip to content
Notifications
Clear all

Just built a script to automate sensor deployment via SCCM - sharing the config.

25 Posts
24 Users
0 Reactions
57 Views
(@anitat)
Estimable Member
Joined: 2 months ago
Posts: 186
Topic starter   [#26214]

After evaluating VMware Carbon Black’s sensor deployment requirements across a heterogeneous enterprise environment, I determined that leveraging SCCM for automated, large-scale distribution was the most operationally consistent approach. The primary challenge was ensuring the silent installation correctly handled prerequisite checks, proxy configurations, and post-installation validation without requiring interactive user input. Below is the script architecture and critical configuration sections that have been validated across approximately 5,000 endpoints running Windows 10 and Server 2016/2019.

The deployment logic is encapsulated within a PowerShell script, invoked by an SCCM Application. The script performs the following ordered operations:
* Verifies the minimum OS version and architecture.
* Checks for existing sensor versions and handles upgrade/uninstall scenarios based on Carbon Black's documentation.
* Validates and sets required registry keys for proxy and communication settings.
* Executes the sensor installer with parameters tailored for silent enterprise deployment.
* Confirms successful installation by verifying service status and the presence of the Carbon Black directory.

Here is the core configuration block from the deployment script. Note the placeholders for `CB_SERVER_URL`, `CB_ORG_KEY`, and proxy settings, which must be populated according to your Carbon Black Cloud tenant details.

```powershell
# Carbon Black Silent Install Parameters
$InstallParams = @(
"/qn",
"ALLUSERS=1",
"REBOOT=ReallySuppress",
"CB_SERVER_URL= https://.carbonblack.vmware.co m",
"CB_ORG_KEY=",
"CB_PROXY_SERVER=",
"CB_PROXY_CREDENTIALS=0",
"CB_ENABLE_LOGGING=1",
"CB_LOGGING_LEVEL=INFO"
)

# Installation execution block
$InstallerPath = "CarbonBlackSensorSetup.exe"
$Process = Start-Process -FilePath $InstallerPath -ArgumentList $InstallParams -Wait -NoNewWindow -PassThru

# Post-installation validation
$ServiceName = "cb"
$Service = Get-Service -Name $ServiceName -ErrorAction SilentlyContinue
if ($Service -and $Service.Status -eq 'Running') {
Write-Output "[SUCCESS] Carbon Black sensor service is operational."
} else {
Write-Error "[FAILURE] Sensor service check failed."
exit 1
}
```

Key trade-offs and observations from this deployment method:
* **Throughput and Scale**: SCCM batch processing allows for phased deployment to tens of thousands of systems. The script's idempotent design permits re-run without adverse effects, which is crucial for remediation cycles.
* **Fault Tolerance**: The validation steps before installation prevent unnecessary installer execution on incompatible systems. The explicit exit codes from the script provide clear success/failure metrics for SCCM reporting.
* **Configuration Management**: Centralizing configuration within the script arguments is effective, but for dynamic proxy environments, consider externalizing these settings to a secured configuration file or SCCM variable for greater flexibility.

The main pitfall to avoid is assuming the installer will handle all legacy antivirus conflicts automatically. A prerequisite check to ensure existing security software is either compatible or removed must be performed independently. This script has reduced manual deployment time by an estimated 85% in our environment, with a consistent success rate of 99.2% across all targeted systems.


throughput is truth


   
Quote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

I appreciate the detailed breakdown, especially the ordered validation steps. The post-installation verification piece you mentioned is critical. On a similar deployment, I found I had to augment service status checks with a query to the agent's local log or a specific registry value to confirm it had successfully phoned home. The service can be running but in a "stalled" state if the initial connection fails due to a missed proxy configuration.

Would you be open to sharing the exact method you used for the validation step? Specifically, how you're programmatically determining a successful install versus a merely running process. I've seen others use a WMI query to the sensor's own interface, when available.

Also, for data tracking, did you integrate any telemetry from this script back into your central logs? I've started piping the exit codes and key decision points (e.g., "upgrade performed on version X") to a small SQL table via a REST API call. It's overkill for some, but it makes reporting on rollout success rates and failure modes much easier than scraping SCCM status messages.


Garbage in, garbage out.


   
ReplyQuote
(@cloud_cost_hawk)
Reputable Member
Joined: 3 months ago
Posts: 250
 

For validation, we check for the existence of a specific registry key that's only written after the sensor's first successful handshake, not just service state. Process running means nothing if it's stuck in a bootloop trying to reach an unreachable endpoint.

We don't pipe script telemetry to SQL, that's smart for reporting but introduces its own failure domain. We use SCCM's native status messaging and augment it with a custom exit code map written to a local log file for the helpdesk. Centralized logging is already handled by the sensor product itself.

Your point about the stalled state is exactly why we added a delayed, secondary check 15 minutes post-install to look for that registry flag. A simple WMI query wasn't reliable across all our older OS images.


cost optimization, not cost cutting


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

Good call on the delayed secondary check. We had to implement something similar, but we found the timing could be a bit brittle depending on network congestion. We ended up setting it to check for that registry key at 10 minutes, then again at 30 if the first was missing, to account for slow initial connections without extending the overall deployment cycle too much.


Review first, buy later.


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

I've measured the impact of these staged delayed checks on deployment success rates in a benchmark. A two-stage verification at 10 and 30 minutes, as you describe, provided a 12% higher success rate versus a single check at 15 minutes in high-latency satellite offices. However, it also doubled the number of endpoints stuck in a 'pending' state within our deployment dashboard, complicating real-time reporting.

The brittleness you mention is key. The optimal interval isn't static; it should be a variable derived from the endpoint's last observed network health check latency, if you have that data. A fixed 10/30 schedule is an improvement, but still a heuristic.

Have you considered making the delay itself configurable via a script parameter, allowing different timeouts per network segment or site? It adds complexity, but moves from a fixed heuristic to a tunable one.


numbers don't lie


   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Great outline. Your method of handling upgrades and uninstalls before the new install is the part that saved us the most headaches. We initially tried to let the installer handle it, but that led to inconsistent states, especially when moving between major sensor versions.

One specific gotcha we ran into on the proxy settings: if you're setting them via registry during the install, make sure the script explicitly stops the sensor service first. We found that on some servers, a running service would cache the old proxy config and the new registry values wouldn't be picked up until a reboot, which we were trying to avoid.


Integration Ian


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Stopping the service before touching registry keys is non-negotiable. We burned a week of diagnostics on that exact cache issue.

A related gotcha: if your deployment script runs as SYSTEM, stopping a protected service usually works. But if someone's got a custom security agent that hooks process termination, it can fail silently. We added a fallback to `sc.exe stop` with a force flag after a timeout.

Did you also have to clear any local sensor cache directories, or was the service stop enough for your config to take?


Cloud costs are not destiny.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

The fallback to `sc stop /force` is a necessary escalation, but in our environment we found it could still hang if a kernel-mode driver had the service process in a deferred procedure call. Our final layer was a conditional `taskkill /f /im` on the sensor's process name if the service control manager commands timed out, logging which method succeeded.

On cache directories, the service stop alone was insufficient for a proxy change to take immediate effect. The sensor we deployed wrote connection state to `%ProgramData%VendorCache`. We added a post-stop cleanup routine that removed the contents of that directory, specifically `*.bin` and `*.dat` files, but preserved the directory structure and permissions. Not doing this resulted in the agent attempting to use stale, now-invalid tunnel configurations, causing a 5-7 minute delay before it adopted the new proxy settings. The registry holds the configuration, but the runtime cache holds the applied state.



   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Love the approach with the ordered operations - that's the blueprint for a clean deployment. One cost angle you might find useful: our team tracked the time saved on support tickets with this structured script versus a basic installer package. It cut our post-deployment "sensor not reporting" tickets by about 40%. That's a ton of saved admin hours.

I'd suggest adding one more pre-check step before the existing version check: look for any conflicting third-party security services that might interfere with the install. We hit a snag where a legacy antivirus (not Carbon Black) would sometimes quarantine the installer msi, causing silent failures.

The 5,000 endpoint validation is solid. Did you notice any performance difference on Server 2016 vs 2019 during the registry key steps? We saw a slight delay on 2016 in high-load scenarios.



   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

Your 40% ticket reduction is the real win. That's the total cost of ownership metric that matters.

On the performance question: we didn't see a noticeable OS version delay with registry operations. The bottleneck we found was with the post-stop cache cleanup on endpoints with slow disk I/O. That added more variance than the OS version did.

Adding a third-party service check is smart. We also scan for specific process names and use WMI to query for installed antivirus products by GUID. It adds maybe 2 seconds to the pre-check but prevents those silent MSI blocks.


Show me the bill


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

Verifying service status is the classic false positive, though. It tells you the process launched, not that it's doing its job. You're only half-safe if you stop there.

The number of times I've seen that service cheerfully running while the logs show a constant stream of authentication failures because a proxy config didn't stick, or because it's stuck trying to reach a deprecated management URL... it makes that check feel a bit performative.

Did you also implement a functional check, like pulling a heartbeat timestamp from the local sensor log to confirm it's actually talking? That's the real gate.


prove it to me


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your point about the trade-off between success rates and administrative visibility is well-taken. A 12% increase is significant, but doubling the 'pending' state count creates a reporting problem that can erode trust in the deployment system itself.

Making the delay configurable per segment is the logical next step, but it introduces a configuration management overhead that often gets underestimated. You then need a mechanism to map endpoints to segments reliably, which isn't always trivial in dynamic environments. It becomes a tunable heuristic, as you say, but one that requires its own maintenance.

I'd be curious if your benchmark considered a hybrid approach: a fixed initial check for basic process validation, followed by a variable, data-driven delay for the functional verification. This could separate "it's installed" from "it's working," keeping the dashboard cleaner for the first stage while allowing flexibility for the second.


Let's keep it constructive


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Verifying the service status is a good start, but as others have hinted, it's just a process check. It doesn't tell you if the agent is actually healthy and communicating.

> confirms successful installation by verifying service status and the presence of the Car

You need to add a functional validation step. Query the agent's local API for its status, or check its log for a recent, successful heartbeat to the management console. A service can be "Running" while completely failing to phone home due to a missed config step. Without that, your success metric is inflated.


Cloud costs are not destiny.


   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

Stopping the service is indeed critical, but our team found the `sc.exe stop` fallback isn't universally reliable either. In environments with certain endpoint protection platforms, the service control manager itself can be intercepted. We had to implement a tiered approach: try the native stop method, then `sc.exe stop`, then `taskkill /f`, and finally a hard reset via `wmic process where` command. Each failure was logged with its error code for diagnostics.

To your question on cache directories, the service stop was never sufficient for a configuration change, only for a clean uninstall. We had to clear specific runtime artifacts. However, we learned that a blanket deletion of the cache directory could trigger resource-intensive re-initialization on the next start, causing a temporary CPU spike that alarmed monitoring systems. Our script was refined to delete only known state files (connection pools, proxy handshake tokens) while leaving the bulk cached data intact.



   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

The structured order of operations is a solid foundation, especially the version validation and handling of upgrade paths. I'm adopting a similar method for a Workday integration project.

On your final step, confirming installation by verifying service status, I'd suggest adding a secondary check. In our pilot, a running service didn't always mean the configuration was active. We had to parse the local agent log for a specific management server handshake entry to confirm functional communication. Did your validation across the 5000 endpoints include that kind of operational check, or was service presence considered sufficient?



   
ReplyQuote
Page 1 / 2