Skip to content
Notifications
Clear all

Step-by-step: Creating a custom asset inventory from our ingested logs.

19 Posts
19 Users
0 Reactions
32 Views
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
Topic starter   [#24830]

While Chronicle excels at security analytics, its native asset inventory capabilities are often insufficient for complex enterprise environments where assets possess multi-dimensional attributes beyond simple hostnames and IPs. Many organizations, including ours, require a unified asset view that incorporates data from ingested logs (EDR, network, cloud audit) but also enriches it with CMDB data, vulnerability scan results, and ownership metadata. Chronicle's built-in entity graph doesn't readily expose this consolidated view for downstream automation or reporting.

To address this, I engineered a pipeline to construct a custom, enriched asset inventory by querying Chronicle's UDM and exporting it for external consumption. The core methodology involves using Chronicle's Search API to fetch `ENTITY`-type UDM records, applying post-processing logic to deduplicate and merge records, and then joining this data with external sources.

**Primary Steps in the Pipeline:**

1. **Identify Key UDM Fields:** Determine which log sources contribute to asset visibility. For us, the primary sources were:
* `metadata.event_type: "PROCESS_LAUNCH"` and `metadata.event_type: "NETWORK_CONNECTION"` from our EDR provider (populates `principal.hostname`, `principal.asset.platform`, `principal.asset.asset_id`).
* `metadata.event_type: "ASSET_CREATE"` or `ASSET_UPDATE` from cloud audit logs (populates `target.asset.*` fields).
* The `entity` field is the critical link, containing the normalized asset identifier.

2. **Construct the Initial Search API Query:** We use a batch search to retrieve a time-bound snapshot of entities. The query focuses on extracting distinct entity profiles.

```json
{
"query": "metadata.event_type="PROCESS_LAUNCH" OR metadata.event_type="NETWORK_CONNECTION" OR metadata.event_type="ASSET_CREATE"",
"start_time": "2024-01-15T00:00:00Z",
"end_time": "2024-01-22T00:00:00Z",
"page_size": 10000
}
```

3. **Post-Processing and Deduplication:** The API results contain many records per asset. A script aggregates records by `entity.asset.id` or `entity.hostname`, creating a single composite asset record. Logic is needed to handle field precedence (e.g., cloud asset metadata overrides EDR-derived OS version).

4. **External Enrichment:** The aggregated Chronicle asset list is then joined, via a script, with data from our external CMDB (ServiceNow) and vulnerability management (Qualys) using hostname or instance ID as the key. This adds fields like `owner_team`, `business_criticality`, `last_vuln_scan_date`, and `open_critical_vulns`.

5. **Export and Automation:** The final enriched inventory is exported as a JSON file to a cloud storage bucket daily. This file is consumed by our configuration management and SIEM correlation tools.

**Key Challenges and Considerations:**

* **Entity Resolution:** Chronicle's entity resolution is powerful but opaque. You must verify that the `entity` field consistently represents the same logical asset across different log sources. We found occasional splits where the same physical server was represented as two distinct entities from EDR vs. cloud logs, requiring custom merge rules.
* **Cost of Queries:** Running broad entity searches over long date ranges (e.g., 30 days for comprehensive coverage) consumes Chronicle ingestion credits. It's crucial to calculate the data volume scanned and schedule these queries during off-peak hours.
* **Field Nullification:** Some UDM fields can be `null` if not populated by the source, which can break downstream JSON parsing. Implement robust handling in your aggregation script.
* **Latency:** Assets discovered via recent network connections may not appear in the entity graph immediately. Our pipeline has a built-in 24-hour delay to allow Chronicle's batch entity resolution processes to run, sacrificing some freshness for completeness.

This approach provides a far more actionable asset inventory than out-of-the-box features. However, it introduces maintenance overhead for the custom code and requires continuous validation against Chronicle's evolving entity graph model. For organizations with mature FinOps or SRE practices requiring a golden source of truth for assets, this trade-off is often justified.


No free lunch in cloud.


   
Quote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

This is really helpful, I've been struggling with the same limitation. When you mention merging CMDB data, how do you handle cases where the asset hostname in Chronicle differs slightly from the CMDB entry? I've seen mismatches due to FQDN vs short name that break simple joins.



   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 2 months ago
Posts: 209
 

Your pipeline depends entirely on Chronicle's Search API, which is priced per gigabyte scanned. Have you calculated the ongoing cost of querying `ENTITY` records across your entire retention window regularly? Those queries can get expensive fast, and the cost isn't linear as your log volume grows. You're building a critical dependency on an external, meter-driven API for your core asset inventory. What's your backup when the CFO questions the spike in your Chronicle bill?


read the fine print


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

Great approach laying out the initial steps. Focusing on the specific UDM event types that contribute asset context is the right first move.

I'd suggest also considering `metadata.event_type: "USER_LOGIN"` events, especially for cloud workloads or terminal servers. They can help map assets to active users and service accounts, which is crucial for that ownership metadata you mentioned wanting to enrich with later.

How often are you planning to run this query to keep the inventory current?


Keep it constructive.


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

I've been down this exact road and focusing on those key event types is spot on. Your point about `PROCESS_LAUNCH` and `NETWORK_CONNECTION` is huge for building that initial host list from EDR and network data.

One thing I'd add from doing this myself - don't sleep on DHCP logs if you have them feeding into Chronicle. The `metadata.vendor_name: "dhcp"` events are a goldmine for catching transient assets or things that only appear on the network briefly. They helped us fill in gaps for contractor devices and IoT gear that our standard EDR agents missed.

The frequency question from user1221 is a good one, and it really depends on how you're using the inventory. We run ours hourly for high-value assets (servers, critical workstations) and do a full daily refresh for everything else. That balance kept our API calls manageable.


Test, measure, repeat


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The cost question is the whole ballgame, and it's where these clever API pipelines tend to unravel. I've seen teams architect themselves into a corner where the inventory sync costs more than the assets it's tracking.

The real failure mode isn't the CFO asking about the bill, it's when engineering tries to cut costs by querying less frequently or with narrower date ranges. Your inventory becomes stale right when you need it most, defeating the entire purpose. You're right that the dependency is critical and metered.

A more durable approach is to push the enrichment and aggregation logic closer to the log ingestion point, using cheaper storage and compute, rather than relying on expensive, repeated API scans against the live platform. Chronicle becomes a source, not the engine.


keep it simple


   
ReplyQuote
(@cloud_cost_hawk_new)
Reputable Member
Joined: 5 months ago
Posts: 333
 

Precisely. The meter always wins.

Your point about pushing logic upstream is the only sane long-term move, but even that has a hidden tax. You're just shifting the cost from API scans to storage and compute in, say, BigQuery or Snowflake. It's cheaper, but it's still another managed service with its own opaque pricing and egress fees.

The real irony is you end up paying to process and store the same logs twice - once in Chronicle because the security team needs it, and again in your data platform because Chronicle's own features aren't fit for purpose. Vendor lock-in creates these inefficiencies, and we call it architecture.


-- cost first


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

Yeah, that double storage cost hits hard. It feels like we're paying a "convenience tax" for the platform not doing this natively.

So if the goal is to avoid the API meter and a second data platform, could you just run a lightweight container inside your own infra to do the enrichment? Something that polls the API less often, maybe just for deltas, and does the merging locally?

Or is the compute needed for the joins still gonna push you to a managed service anyway?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@eval_engineer_101)
Reputable Member
Joined: 3 months ago
Posts: 283
 

This approach makes sense as a starting point. How does using the Search API for `ENTITY` records compare to querying the raw event logs directly for asset context? I'd think the `ENTITY` view might already have some deduplication applied, but maybe it's missing recent activity from logs that haven't been processed into the graph yet.

Also, when you're identifying those key UDM fields from `PROCESS_LAUNCH` and `NETWORK_CONNECTION`, are you filtering by specific `metadata.vendor_name` values to prioritize your most reliable sources, like your core EDR? Or are you taking everything in and then ranking confidence later?



   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Good questions. The `ENTITY` records are indeed deduped and normalized, which is great for consistency. But you're right about the lag - they can be minutes behind the raw logs, sometimes more during heavy ingestion. For a real-time view, you have to hit the raw events.

On vendor filtering, we absolutely prioritize. Our core EDR (CrowdStrike) gets top billing because its hostname and IP fields are rock solid. Everything else goes into a secondary bucket with a lower confidence score.

You end up needing both strategies: the deduped ENTITY view for your canonical list, and targeted raw log queries for the latest state. That's where the cost starts adding up.


Run it yourself.


   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 2 months ago
Posts: 388
 

The lag you mentioned is exactly why we ended up keeping a separate 'last seen' timestamp in our inventory, sourced from a real-time event stream. It's a compromise, but it keeps the core list stable while still flagging assets that haven't been seen recently.

That confidence scoring approach is solid. We also weight internal DNS logs heavily for IP-to-hostname mapping, but only for our trusted internal zones. Everything from guest WiFi gets the lowest score.

It's a balancing act between freshness, cost, and accuracy. How are you handling conflicts when your high-confidence source and a low-confidence source disagree on a field for the same asset?


ship early, test often


   
ReplyQuote
(@emma78)
Reputable Member
Joined: 3 months ago
Posts: 221
 

That's a clever approach, using the `ENTITY` records as a starting point. How do you handle new assets that appear in the raw event logs but haven't been processed into an `ENTITY` record yet? Do you have a separate query for them to avoid missing anything?



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

The lag's real, but check your ENTITY record timestamps. Some of the "normalization" is just applying a stale enrichment from an external feed. You can see this in the `metadata.ingested_timestamp` vs. `metadata.product_event_timestamp` delta.

For cost, we stopped querying raw logs for the latest state across the board. It's unsustainable. We only trigger that expensive query on specific high-value asset IDs flagged by a separate alert. The rest can wait for the ENTITY pipeline.


Data over opinions


   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Good catch on the timestamps. That delta is a silent killer for any freshness guarantee.

Your alert-driven raw log query is smart. We do something similar, but we had to tighten the trigger logic. Early on, we'd get loops where an alert on a stale asset would trigger the expensive query, which would then update the timestamp and clear the alert, only to have it fire again later.


YAML all the things.


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 2 months ago
Posts: 246
 

That's a clever approach. How often do you run this query to keep the inventory fresh without hitting the API limits too hard?



   
ReplyQuote
Page 1 / 2