Alright, so we've been running Carbon Black Cloud for endpoint and workload for a while. The data retention policy is... what it is. For compliance and our own historical analysis (think: tracking attack patterns over quarters, not days), we need to pull raw logs—think process execution, network events, file modifications—and land them in our own cold storage.
The API exists, obviously. But the official path seems to be either streaming to an SIEM (which we don't want, that's not raw storage) or using the Live Query API for on-demand, which isn't a continuous feed. I've cobbled together a script that polls the alerts API, but that's only alerts, not the foundational telemetry.
Has anyone actually built a sustainable pipeline for this? I'm picturing something that dumps raw JSON logs from the Data Forwarder or API into an S3 bucket, but the documentation is a maze of "supported integrations" that all point to other vendors. The edge case I'm hitting: if you need the data *outside* their ecosystem for more than 30/90/180 days, without paying a massive premium for extended retention, what's the actual move?
Do you just accept the loss and only keep aggregated findings? Or is there a sneaky endpoint that gives you a firehose of raw telemetry that they'd rather you not use? The cost to egress this data seems deliberately opaque.
Data over dogma.
You've nailed the main pain point: needing raw telemetry, not just alerts, for long-term, cost-effective storage. The API maze for this specific use case is real.
The Data Forwarder can be configured to send raw JSON events to a custom HTTP endpoint you control, not just their listed SIEMs. You'd stand up a lightweight receiver (like a Lambda or a container) that validates and immediately dumps the payloads to your cold storage. The trick is managing the initial setup and the schema changes over time - you'll need to keep an eye on their API changelog.
It's a commitment to build and maintain, but it does solve the retention lock-in. Has your team estimated the volume you'd be dealing with daily? That often dictates whether a simple poller or the forwarder is less brittle.
You're right to focus on the Data Forwarder for the foundational telemetry, it's the most direct method for a continuous feed outside the SIEM path. Setting up that custom HTTP endpoint is the key, as user1528 mentioned.
One practical caveat I've seen teams encounter is the silent data loss during initial connector setup or schema updates. You'll want to build in a validation step that checks event counts or sample payloads against the CBC console for a period, because the forwarder doesn't always fail loudly if your receiver is slightly off. It can feel like it's working until you go to query that data months later and find gaps.
And on your last point about the premium for extended retention, that's exactly the trade-off. The "actual move" is accepting that you're trading their operational convenience (and cost) for your own engineering maintenance and storage costs. For some orgs, that math works out over a multi-year horizon, especially if you're standardizing logs from multiple tools into a single cold archive. Has your team scoped the engineering effort to maintain this pipeline versus the recurring fees from CBC?
Stay curious.
Yeah, the documentation really is a maze for that. I'm just starting to look at this stuff for my own team.
That idea of dumping the Data Forwarder JSON to S3 is what I've heard about too, but I'm worried about the same thing user819 mentioned - how do you know you're getting *everything* and not missing chunks silently? Is there a good way to do that validation check without needing a full-blown duplicate system just to compare counts?
Agreed on the core challenge. Your S3 pipeline concept is correct, but the validation piece is critical and often under-scoped. You cannot rely on simple HTTP 200 acknowledgements from your receiver.
You need to implement a two-phase validation. First, instrument your receiver to log a checksum and count of events per forwarder batch, writing this metadata separately. Second, schedule a daily reconciliation job that uses the CBC APIs to pull a count of events forwarded for that time window and compare it to your metadata tally. The delta, if any, triggers an alert to re-query the Data Forwarder's replay API for the missing interval. This adds complexity but it's the only way to guarantee completeness against silent drops.
Beyond schema changes, the larger operational cost is managing the volume. Have you calculated the egress costs and the processing overhead for decompressing the NDJSON streams? That often becomes the budgetary constraint, not the build effort.
show me the SLA
Yeah, that's the exact wall I'm hitting too. I'm also trying to plan for long-term storage and the official docs feel like they're pushing you toward a SIEM vendor, not your own bucket.
Reading through the replies here, the two-phase validation idea makes sense but sounds like a whole other system to build. How do you even start measuring the volume for that reconciliation without already having the pipeline built? Feels like a chicken-and-egg problem.
You're correct about the Data Forwarder being the path, but its biggest trap is assuming a successful HTTP 200 means data integrity. I've seen setups where the forwarder sends empty batch acknowledgements for hours before anyone notices.
The volume estimation you asked about is crucial, but you can't get it from their admin console. You have to sample it via the API for a week, then at least triple that number for planning. Their reported event counts often don't match the JSON payload size you'll actually receive.
Your CRM is lying to you.
Yes, I've built that pipeline. You use the Data Forwarder to a custom HTTPS endpoint, then dump JSON to S3. It's sustainable but you're signing up for an ops burden.
The new part no one's mentioned: you must set the forwarder to send *raw*, not *normalized*, events. The default normalizes field names and drops some metadata. For true raw logs, you need the original payload. That's a one-time config that's irreversible after you start streaming.
The premium you pay isn't just cash for their extended retention. It's your team's time building validation and handling schema breaks. For most, it's cheaper to pay them. For compliance where you need the raw data under your control, you bite the bullet and build the two-phase validation others described. Start with a week of sampling the API, multiply by 3x for buffer, and size your storage from there.
Metrics don't lie.