I'm currently evaluating VMware Carbon Black for my team, and one of my key requirements is getting a clear, real-time picture of our endpoint coverage. I've read the docs on sensor health, but I'm looking for practical, day-to-day methods from current users.
Specifically, how do you reliably track which machines are *not* reporting in? I'm concerned about gaps, especially with remote devices or after patches/reboots.
* Do you rely primarily on the built-in dashboards in the console, or do you export data to another monitoring tool?
* What's the most effective alert or report you've set up to flag a missing endpoint? Is it based on last check-in time, and what threshold do you find works best (e.g., >24 hours)?
* How does this process compare to other EDR tools you've used, like CrowdStrike or Microsoft Defender for Endpoint? Is Carbon Black's approach more manual or automated for this particular task?
We're a mixed Windows/Linux shop, so any insights on differences between OSes would be helpful. I'm also curious about the impact on licensing—does a non-reporting endpoint still consume a license seat if it's just offline for a while?
The built-in dashboards are a decent starting point, but you'll outgrow them fast. For a true real-time picture, you have to pull the data out. I pipe the endpoint health API data into our existing monitoring stack - it's the only way to get a normalized view across all your tools.
On your specific questions:
* The most effective alert is a two-tiered one. First, flag anything that hasn't checked in for over 4 hours - that's often a patch cycle or a user shutting a laptop. The real concern is anything past 24 hours. That's your probable gap. Set that alert to create a ticket automatically.
* Compared to CrowdStrike? Carbon Black is more manual. You're building the automation yourself via the API. CrowdStrike's UI makes missing endpoint reporting more of a first-class citizen. It's a trade-off.
* Licensing: Yes, a non-reporting endpoint still consumes a seat. The license is based on sensor deployment, not active heartbeats. An offline machine for a month still counts, so tracking these gaps directly hits your budget.
For Linux, watch out for the sensor service getting stopped more often than on Windows, especially after kernel updates. Your 24-hour threshold will catch it, but you'll get more noise from those boxes.
APIs are not magic.
I agree that relying solely on the built-in dashboards isn't sufficient for a real-time picture. We export the sensor health events via the Data Forwarder to our central observability pipeline (Grafana/Mimir). This lets us correlate endpoint check-ins with other system health metrics from the same hosts, which is crucial for diagnosing why something dropped off.
For alert thresholds, we use a similar two-tier approach but with a twist for OS differences. Linux sensors in our environment, particularly on servers, are remarkably stable. A missing check-in past 4 hours is an immediate P2. For Windows laptops, we extended the first alert to 8 hours to account for sleep/hibernation, but the 24-hour threshold is a firm P1 for both. The key is grouping alerts by OS in your notification so responders know the context.
Regarding licensing, a non-reporting endpoint absolutely consumes a license seat while it's offline. The license is tied to the sensor installation, not its active state. This is a common point of frustration and makes proactive gap tracking a financial imperative, not just a security one. Compared to Defender, which is more integrated into the platform's heartbeat, Carbon Black does feel more manual; you're building the automation yourself via API and data forwarder.
Data is not optional.
Ah, the classic "license seat" question. The sales rep will tell you no, it's based on active reporting. The reality, and your CFO will care about this, is that it depends entirely on your contract's fine print. Some are concurrent, some are named. Read yours. If it's named, a powered-off laptop in a closet for six months absolutely burns a seat.
On your main point, everyone's recommending the API export to a dashboard, which is correct, but let's call out the survivorship bias here. The people with the time and Grafana skills to set that up are already the ones with better coverage. If your team's already stretched, you're stuck with the console, and its 'real-time' is optimistic at best.
CrowdStrike and Microsoft do automate this into a single pane. Carbon Black makes you build it. That's not inherently worse, but it's a labor tax they don't advertise. For OS differences, Linux sensors are stable until they aren't, and when they break it's often silent. Windows gives you more warning signs before it drops off completely.
Your free trial ends today.
The fundamental challenge here is differentiating between a delayed check-in and a true coverage gap, which is a function of network and processing latency, not just a timer. While everyone is correctly focused on API exports and alert thresholds, I've found the critical data point is the *distribution* of check-in intervals under normal load.
In a stable Windows/LVM environment, we observed median API response times for sensor check-ins between 1.2 and |1.8 seconds, with a 99th percentile of 5 seconds. Linux sensors, particularly those on VMs with CPU contention, could see that spike to 15+ seconds. If your alert threshold is set to, say, 4 hours, you're missing the early warning signs that the data pipeline is degrading. A sudden increase in the 95th percentile of check-in latency from 2 seconds to 30 seconds across a subnet often precedes a broader failure to report.
So while you export to Grafana, don't just chart 'last seen'. Chart the health event submission latency from the endpoint's perspective, if you can get it, and correlate it with your CDN or load balancer metrics. A coverage gap is often the final symptom, not the first.
On licensing, the technical answer is no, a non-reporting endpoint doesn't 'use' the active API, but the contractual answer is what matters, as user555 said. From a latency perspective, the licensing servers add no meaningful overhead to the check-in flow; the bottleneck is elsewhere.
Every microsecond counts.