A common challenge in modern security operations is maintaining an accurate, queryable asset inventory. Data is often siloed across CMDBs, cloud provider APIs, vulnerability scanners, and EDR agents. Elastic Security, with its underlying Elasticsearch engine, provides a powerful platform to unify this data, but the process of building a coherent inventory from disparate sources requires careful data modeling and ingest pipeline design.
The core concept is to treat each asset as a living document that is updated incrementally by various source streams. The goal is not to replace source systems, but to create a consolidated "golden record" enriched with the most recent and relevant data from each. The primary trade-off is between schema rigidity, which ensures consistency, and schema flexibility, which accommodates heterogeneous data sources.
**Key Data Model Decisions:**
* **Document ID:** Use a persistent, unique asset identifier (e.g., `host.id` or `asset.guid`) as the `_id`. This ensures all updates from different sources for the same asset target the same Elasticsearch document.
* **Update Strategy:** Utilize the `upsert` capability of the Update API. The document structure should define clear fields for each data source.
* **Timestamp Management:** Maintain separate timestamp fields for each data source update (e.g., `last_seen_cmdb`, `last_seen_edr`). A composite `asset.last_updated` field can be derived from the most recent of these.
Here is a simplified example of an asset document structure after enrichment from two sources:
```json
{
"asset": {
"id": "host-12345",
"hostname": "prod-db-01",
"ip_address": ["10.0.1.15", "fe80::..."],
"last_updated": "2023-10-27T10:24:00Z"
},
"source": {
"cmdb": {
"owner": "Database Team",
"cost_center": "CC-750",
"last_updated": "2023-10-26T08:30:00Z"
},
"edr": {
"agent_version": "7.10.1",
"last_seen": "2023-10-27T10:24:00Z",
"os": {
"platform": "windows",
"version": "Server 2022"
}
}
}
}
```
**Implementation Workflow:**
1. **Ingest Pipeline Creation:** Build an ingest pipeline to normalize incoming data. This pipeline should:
* Extract or generate the canonical `asset.id`.
* Map source-specific fields into the nested structure (e.g., `source.cmdb.owner`).
* Set the relevant source timestamp.
* Calculate/update the top-level `asset.last_updated`.
2. **Data Ingestion:** Configure your data shippers (Fleet Integrations, Logstash, or the Elasticsearch API) to target a dedicated asset index (e.g., `assets-current`) and apply the ingest pipeline. The update operation is crucial:
```json
POST assets-current/_update/{{asset.id}}
{
"scripted_upsert": true,
"script": {
"source": """
ctx._source.asset = params.asset;
ctx._source.source.cmdb = params.source.cmdb;
ctx._source.asset.last_updated = ZonedDateTime.parse(params.source.cmdb.last_updated).toInstant().toEpochMilli();
""",
"params": {
"asset": { "id": "host-12345", "hostname": "prod-db-01" },
"source": { "cmdb": { "owner": "Database Team", "last_updated": "2023-10-28T09:15:00Z" } }
}
},
"upsert": {}
}
```
3. **Querying and Visualization:** Once populated, the inventory can be queried in Kibana using standard Elasticsearch queries. For example:
* Find all assets where `source.edr.agent_version` is older than 7.9.0.
* Correlate assets from a specific `source.cmdb.cost_center` with vulnerability data from a joined query on another index.
* Build Lens visualizations showing OS platform distribution or agent coverage gaps.
The main pitfalls of this approach are identity resolution (ensuring `asset.id` is consistent) and data staleness. A companion index holding a heartbeat or regular check-in from agents can help flag assets that have not reported recently, allowing for automated cleanup scripts to mark them as potentially decommissioned. This model's strength lies in its ability to provide a real-time, queryable view of asset state, which is foundational for effective vulnerability management, incident response, and security posture reporting.
brianh
Oh, this is fascinating. It reminds me of building a single customer view in a CRM, trying to merge data from marketing automation, support tickets, and sales calls into one profile. That same "golden record" concept is so critical.
Your point about the primary trade-off really hits home. In marketing, we face that exact same tension: do we enforce a strict schema for clean analytics, or stay flexible to capture unique data points from new channels? We often end up with a core, rigid schema for key identifiers (like your `host.id`) and then a flexible nested object for "everything else" that we can still query, just not as efficiently.
How do you handle versioning or the history of an asset's state? In our world, we sometimes need to know not just the current merged profile, but what a customer's status was *last week* for accurate campaign attribution. Do you use a separate index for historical snapshots, or is that handled within the single document?
If it's not measurable, it's not marketing.
Your point about using the Update API's upsert with a persistent ID is the correct starting point, but in practice that document can become a massive, unmanageable aggregation of every historical update if you aren't careful. We had to implement a companion cleanup process.
We ended up defining a strict schema for the core golden record fields and routing all updates through a painless script in the ingest pipeline. This script would evaluate incoming source data, merge it into the core fields, and move deprecated or versioned data out into a nested history array within the same document. That preserved the queryability of the current state while still keeping a lightweight audit trail inline.
Without that step, our documents ballooned and performance on the inventory index degraded significantly after a few months of updates from ten different source systems.
Latency is a liability
The trade-off you mention between schema rigidity and flexibility is the part that always trips me up. I've used Looker to try and model similar merged data, and if the underlying Elasticsearch schema isn't right, the LookML gets messy fast.
Is there a rule of thumb for which fields belong in the rigid core versus the flexible part? In my cost analytics work, something like `instance_type` feels core, but `current_month_cost` definitely doesn't.
Great CRM analogy, it's the same problem space. We handle history almost exactly as user1545 described in their follow-up, using that nested history array within the document for the audit trail.
Storing a full historical snapshot in a separate index would be more correct for heavy temporal querying, but it's overkill for us. For campaign attribution, would querying a nested history array be too slow on your scale? I'm curious if a time-series index for major state changes would be the next step if we needed to analyze trends.