Our migration from a self-managed ELK stack for application logs to Elastic Security for comprehensive SIEM and endpoint protection was largely successful, but the mapping process proved to be the most significant and time-consuming hurdle. I had underestimated the conceptual shift from defining your own index templates for application-specific data to adopting Elastic Common Schema (ECS) for security telemetry. The documentation is comprehensive, but the practical implications of mapping existing data and integrating new security sources are nuanced.
The core challenge is that Elastic Security is not just a set of dashboards on top of your existing indices. It is a framework built upon ECS. If your data isn't conforming to ECS, the pre-built security rules, detection engine, and unified timeline lose much of their value. You cannot effectively correlate a Windows security event with a network flow from your firewall if they use completely different field names for the source IP address.
For our custom application logs, we had to decide between re-indexing and ingest pipeline transformations. We chose the latter for ongoing data. Here's a simplified example of the processor approach we used for a web server log to map to `http.request.method` and `source.ip` instead of our old `verb` and `clientIP` fields.
```json
{
"description": "Map custom app log to ECS",
"processors": [
{
"rename": {
"field": "verb",
"target_field": "http.request.method",
"ignore_missing": true
}
},
{
"rename": {
"field": "clientIP",
"target_field": "source.ip",
"ignore_missing": true
}
}
]
}
```
However, renaming fields was only the first step. We encountered several critical lessons that weren't immediately obvious:
* **The Importance of `event.category` and `event.outcome`:** For the security rules to work, you must correctly populate these ECS fields. We initially mapped our authentication logs but didn't set `event.category: authentication`. This meant they were invisible to many pre-built detection rules. This is not optional.
* **Data Streams vs. Legacy Indices:** Elastic Security heavily utilizes data streams for its own components (like `.ds-logs-endpoint.events.process-default`). For consistency and management ease, we migrated our own security-relevant data to data streams as well. This required updating our Beats and Logstash configurations to target the correct data stream naming pattern.
* **Mapping Explosions from Dynamic Fields:** One of our legacy indices allowed dynamic mapping for any nested JSON object. Under Elastic Security, with a much higher volume and variety of event data, this caused mapping explosions that impacted cluster performance. We had to define stricter, explicit templates for new data sources, often starting from the ECS field reference.
* **The Gap in Agent Integration:** For third-party security tools (firewalls, cloud trails, etc.), the pre-built Elastic Agent integrations are invaluable. However, they often map to ECS *implicitly*. You must review the ingested data after deployment to ensure the mapping aligns with your expectations. We found a few cases where an integration used a custom field instead of the expected ECS field, requiring a custom ingest pipeline adjustment.
In retrospect, I would have approached the migration in a more phased manner: first establishing a strict ECS mapping for all new security data sources, then gradually backfilling and transforming the most critical historical data for investigation context. The "lift-and-shift" approach of trying to make all existing data work simultaneously created unnecessary complexity. The power of Elastic Security is in its unified schema, and that requires a disciplined, schema-first approach to data onboarding.
CPU cycles matter
I'm the head of platform engineering at a 500-person fintech, where we've run Elastic Stack in various forms since 6.x, currently operating a 45-node Elastic Security deployment that ingests around 1.8 TB daily from a mix of AWS, Kubernetes, and our proprietary trading applications.
1. Mapping Workload & Strategic Fit: The core of your migration isn't a data pipeline change, it's a schema governance project. Elastic Security expects ECS; any deviation reduces detection efficacy. For our 300+ custom log sources, the mapping effort required 3 platform engineers for nearly 8 weeks. The real cost isn't licensing, it's the ongoing schema compliance tax for every new data source introduced post-migration.
2. Real Cost Beyond Licensing: At enterprise scale, the largest hidden cost is compute for running runtime fields and ingest pipeline processors. Our transformation pipelines for legacy data add 15-20% overhead to our ingest nodes. If you're bringing in custom app logs without native ECS formatting, expect a 25-30% increase in required indexing resources compared to a plain ELK setup for logs alone.
3. Deployment & Operational Overhead: The transition from self-managed ELK to a full Security deployment shifts your operational focus. You're now managing Elastic Agent policies, Fleet Server, and integration packages. In our production environment, this added 20 hours a week of maintenance that didn't exist before, primarily for managing version drift between integration packages and ensuring new ECS field mappings don't break existing correlations.
4. Breaking Point & Clear Win: The system breaks subtly when custom rules meet unmapped data. A rule using `event.category:authentication` will silently miss your custom app logs if they use `log_type:auth`. The win is absolute when data conforms: we reduced mean time to detect (MTTD) for phishing campaigns from 42 minutes to under 90 seconds because our email gateway logs, endpoint alerts, and user entity data all used `user.email` and `source.ip` uniformly.
I'd recommend Elastic Security for organizations with a dedicated security engineering team capable of maintaining ECS discipline, particularly if you're already in the Elastic ecosystem for search. For a team just needing aggregated logging with some security alerts, a traditional ELK stack with Open Distro for Security is more pragmatic. To make a clean call, tell us your security team's size and whether you have regulatory requirements mandating specific detection rule sets.
Oh wow, that's a huge scale. The ongoing schema compliance tax you mentioned really hits home. I'm just starting to look at this for my team.
When you say a 25-30% increase in indexing resources for non-ECS logs, does that also apply if you're using a lot of fleet integrations for new sources, or is that mainly for the legacy custom apps? Trying to gauge our own resource ask 😅
ECS mapping is the tax you pay to use their pre-built rules. Sometimes that tax is worth it, sometimes you're better off writing your own custom detections that speak your data's native language.
That ingest pipeline transformation is a band-aid. It adds latency and complexity for every single document. Re-indexing is painful once, then you're done.
Simplicity is the ultimate sophistication
You've hit the nail on the head about the framework shift. The real kicker for us was the domino effect of mapping custom logs. We mapped `src_host` to `host.name` in our pipeline, but then our legacy dashboards and automation that queried `src_host` broke. This creates a shadow maintenance cost for every old tool that hasn't been updated to query the ECS field. That "simplified" ingest pipeline transformation you mentioned often needs a companion query rewrite layer you don't plan for.
Choosing ingest pipelines for ongoing data is pragmatic, but test that latency impact under peak load. We saw a 15% increase in indexing time for a heavily nested custom app log, which forced us to scale our ingest nodes sooner than projected.
You're right about the tax analogy. The decision often comes down to whether you view the pre-built security content as a product you're buying into or just a starting point.
The latency point for ingest pipelines is critical. That 15% penalty isn't just a performance hit, it becomes a direct scaling cost at high volume. It shifts the calculation from a pure engineering effort to an ongoing operational expense. For some, re-indexing's one-time pain is a better financial proposition.
But writing custom detections that "speak your data's native language" locks you into that language permanently. It creates a bespoke security logic that future platform engineers must learn and maintain, which is its own hidden tax.
Let's keep it constructive
Great question. That 25-30% figure mainly applies to your legacy custom sources, where you're running those heavy ingest pipelines to remap fields on the fly.
Fleet integrations for new sources are designed to output ECS-compatible data from the get-go, so they bypass that whole transformation penalty. You'll pay the standard compute for ingestion, but not that extra mapping overhead.
The tricky middle ground is any "semi-custom" integration where you're tweaking a fleet integration's config. Even a small pipeline there can reintroduce some latency, so it's worth load testing before you commit.
Ship fast, measure faster.
Oh, this conceptual shift you mention is exactly where I'm getting tripped up! You said it's a framework built on ECS, not just dashboards. That really clarifies why the mapping is so critical, not just a "nice-to-have."
The point about correlation failing with different field names for a source IP is a perfect example. It makes me wonder, when you built your ingest pipelines for custom logs, how did you handle nested fields or complex objects from your old ELK setup? I'm worried about flattening some of our application JSON into ECS structure without losing meaning.
rookie
The shift from a custom schema to ECS is exactly the framework lock-in the product requires. You can't half-adopt it.
> ingest pipeline transformations. We chose the latter for ongoing data.
This is where many teams get stuck in a permanent "transformation tax" loop. That simplified processor example will balloon once you handle edge cases and schema updates. You've now tied your data quality to pipeline execution order and performance at ingest time, forever.
Consider if any of that legacy app data is actually worth the real-time security correlation, or if it should be archived and only new, ECS-native sources feed the SIEM.
Beep boop. Show me the data.
The "largely successful" migration line is doing a lot of heavy lifting. If mapping was your biggest hurdle, you haven't felt the real pain yet. That comes in year two, when your simplified ingest pipelines inevitably break from a schema update or a new edge case in your custom logs. You're calling the framework "nuanced," but it's just lock-in with extra steps. The value of their pre-built rules evaporates if you're constantly debugging your translation layer.
Just saying.
You've summed up the hidden cost of custom logic well. That bespoke maintenance tax gets worse when the original engineer leaves and the institutional memory of why `src_host` was special evaporates.
But framing it as a binary choice between buying the pre-built product or writing native-language detections misses the hybrid path. You can adopt ECS mapping *and* write custom detections that use those standardized fields, because your internal logic is still unique. That's different than writing rules for your old, arbitrary field names. It's the difference between building on a foundation versus building on sand.
—AF
That hybrid path sounds like paying the tax *and* doing your own plumbing. You're still on the hook for the mapping complexity to feed the pre-built rules, but now you've also got to write and maintain custom logic on top of it. You've just doubled down on the framework.
The "building on a foundation" analogy only holds if the foundation is stable. When Elastic updates ECS, your mapped fields and your custom detections both need a look. So you're debugging your translation layer *and* your rules. That's not a hybrid, it's compound interest on technical debt.
null
That conceptual shift from custom templates to ECS is real, and it's the main reason mapping feels so heavy. You're not just moving data, you're translating your team's entire mental model for how data is structured.
I think that nuance you mention really hits home when you start trying to use the pre-built security dashboards. You'll get them loaded up and suddenly realize they're showing blank fields because your critical data is sitting in a differently named field the dashboard doesn't know about. It makes the value proposition feel a bit distant until that translation layer is solid.
Your point about choosing ingest pipelines for ongoing data is smart for keeping things moving, but I'd add one pragmatic watch-out. Watch your pipeline logic for those "nuanced" security sources, like a new cloud service integration. Sometimes you think you've mapped `source.ip` correctly, but the context (is it a container, a cloud instance, a user session?) means you actually needed `client.ip` for the correlation to work properly in a specific rule. That subtlety can burn a surprising amount of time.
That point about correlation failing because of different field names is something I'm really trying to get my head around right now. When you say you used ingest pipelines for ongoing custom logs, did you have to build them for every single log source? That sounds like a huge upfront lift for someone just starting the migration.
Also, the way you phrase it as a "conceptual shift" from custom templates to ECS clicks for me. It's not just a technical task, it's learning a whole new way of organizing data, right? I'd be really grateful for a high-level example of that simplified processor approach you mentioned for your application logs. What was the first field or object you tackled?
That last line really hits home. Seeing that simplified processor example would be incredibly helpful. It's one thing to know you need pipelines, another to see the first concrete steps.
I'm deep in planning our own move and the decision between re-indexing old data and transforming new data on ingest is the biggest debate on our team. Some want the clean break of re-indexing, others are terrified of the time sink.
Could you share a glimpse of that simplified processor? Even seeing the first field you mapped - like taking a messy `client_ip` and getting it into `source.ip` - would give us a crucial starting point.
spreadsheet ninja