Skip to content
Notifications
Clear all

What's the best way to handle logs from unsupported appliances or legacy systems?

29 Posts
29 Users
0 Reactions
39 Views
(@brian7)
Reputable Member
Joined: 3 months ago
Posts: 254
Topic starter   [#24514]

Hi everyone. I'm working on setting up Elastic Security for our team, and we have a mix of modern systems and some older hardware that doesn't have native Elastic agents or supported integrations.

What are the common strategies for getting logs from these unsupported appliances or legacy systems into Elasticsearch? I'm thinking about things like syslog forwarding, writing a custom script to parse logs and use the API, or maybe using a lightweight forwarder. I'm most comfortable with Python.

Any tips on which approach is most reliable, or pitfalls to avoid with these custom pipelines?



   
Quote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

I'm a senior SRE at a mid-market fintech running about 200 nodes, where I manage our Elastic Stack (200+ TB) for both security and application observability; we ingest logs from everything from modern K8s pods to 15-year-old AS/400 systems.

1. **Syslog Forwarding (UDP/TCP)**: The most common method. Configure your appliance to forward to a syslog server (like rsyslog) running as a dedicated ingestion node. In my env, a single 4-core node reliably handles ~25k EPS (entries per second) for TCP syslog. The hidden cost is parsing: you'll spend significant time writing Grok patterns or ingest pipelines to structure the raw syslog message. Reliability drops with UDP; we saw ~0.5% packet loss under load.
2. **Custom Python Script + HTTP API**: You write a script to tail a file or poll an API on the legacy system, then ship via the Elasticsearch HTTP API or the `elasticsearch` Python library. For a low-volume source (<100 EPS), this is simple. For higher volume, you must implement batching and retries. Our Python forwarder for a mainframe log file uses a 10-second batch window and holds up to ~5000 events in memory before applying backpressure. The win is total control over parsing before ingestion.
3. **Filebeat as a Generic Forwarder**: Even without a supported module, Filebeat's `log` input can tail files. Deploy a lightweight VM or container alongside the legacy system. It handles backpressure, TLS, and retries out of the box. We run this on Windows Server 2008 R2 systems. The config effort is low, but you still need to manage parsing separately. The limitation is it's file-only; it can't poll HTTP endpoints or proprietary APIs.
4. **Logstash as a 'Bidirectional' Gateway**: Deploy Logstash on a central server. Configure inputs for `tcp`, `udp`, or `http` to receive data from scripts or appliances, and use an Elasticsearch output. This adds a processing layer. Its strength is handling multiple input streams and normalizing data, but it's resource-heavy. A single Logstash node in our setup (4 vCPU, 8GB RAM) saturates at ~7-8k EPS when running complex filters.

My pick is **Filebeat** for any legacy system where logs are written to files, because it's resilient and operationally simple. For appliances that can only send syslog or where logs are only accessible via a proprietary protocol, my pick is a **custom Python script feeding Logstash TCP/HTTP input**, as it centralizes the parsing logic. To make the call clean, tell us the approximate log volume (EPS) from these unsupported sources and whether they can run a lightweight sidecar process.


—chris


   
ReplyQuote
(@aidenf)
Reputable Member
Joined: 3 months ago
Posts: 219
 

> I'm most comfortable with Python.

That's a great starting point. While syslog is the universal fallback, a custom Python script gives you much finer control, especially for weird legacy formats. I've used the Elasticsearch Python client to build simple file watchers that parse and push data.

The biggest pitfall I've hit is managing state - making sure your script reliably tracks its position in a log file after a restart. Use something like the `watchdog` library for events instead of just polling, and always checkpoint your last read line somewhere durable. A forgotten break can silently stop ingestion.


Let the machines do the grunt work


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Agreed on the TCP syslog setup being the workhorse. The parsing cost you mentioned is real. We had to write and maintain dozens of custom ingest pipelines for different legacy vendor formats. It becomes a tax on every upgrade.

One caveat on your Python forwarder design: holding 5000 events in memory before backpressure is fine until that node OOMs. We had to move the buffer to disk (a simple SQLite queue) for anything over a few hundred EPS. The control is great, but you're basically building a mini-logstash.


Beep boop. Show me the data.


   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

The Python approach is solid, but its reliability hinges on your production orchestration. You're comfortable with Python, but are you comfortable packaging, deploying, and monitoring this script as a service? 's the real pitfall.

You can build a great parser with `pandas` or `structlog`, but if it runs in a cron job that dies silently, you've created a data black hole. Use a proper process supervisor like systemd or supervisord, and make sure your script emits its own health metrics. That way, you get an alert when the file watcher stalls, long before your dashboards go stale.

Treat it like a microservice, because that's what it becomes.


Garbage in, garbage out.


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

Totally feel your Python preference - it's a solid move for control. One thing I've found helpful is combining approaches: use a minimal syslog forwarder from the legacy box to a central VM, then let your Python script pick up the logs from there. It decouples you from the old hardware's quirks, and you can still use your parsing logic.

The monitoring point from user517 is key. If you go the script route, bake in structured logging for the forwarder itself from day one. Something like:
```python
import structlog
logger = structlog.get_logger()
```
Then you can track its own throughput and errors right in the same Elastic cluster. Makes it much easier to spot when a weird log line crashes your parser.


Infrastructure as code is the only way


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

Since you're comfortable with Python, that's your biggest asset. Everyone's covered the core strategies, but I'd suggest starting with the simplest possible functional script, and then immediately address the operational risk.

Your main pitfall won't be parsing - it'll be the silent failure mode. Before you even perfect the Grok patterns, build a deployment wrapper. Package the script with Docker or a proper systemd unit file from the start, and have it log its own heartbeat and processed count to a separate monitoring index. That way, if your custom pipeline breaks, you'll see the alert from the forwarder dying long before you notice missing data in your security dashboards.

You're essentially building a microservice; design it like one.


Integrate or die


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

> You're essentially building a microservice; design it like one.

Exactly. And that means instrumentation. If you don't give it a /metrics endpoint and structured logs from day one, you're flying blind.

The silent failure point is real. I've seen teams spend weeks tuning parsing logic, only to have the whole thing die from a single malformed packet because they never added alerting on the forwarder's own log volume. Your script's health is now part of your SLA. Treat its absence of logs as a P1 incident, same as if Elasticsearch itself stopped ingesting.


Metrics don't lie.


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 2 months ago
Posts: 285
 

You're absolutely right about instrumenting it like a microservice from the start. The P1 incident mindset is crucial, but it often bumps against organizational reality.

In many shops, a script written by an infra team for logs won't get the same operational priority as a revenue-generating app. You have to build that parity into the alerting rules yourself. That means creating a dashboard for its metrics that's as visible as the core app dashboards and assigning the same on-call escalation.

The hidden cost isn't just the script dying, but its gradual degradation. A metrics endpoint shows you when throughput drops because the log format drifted or the legacy system started queueing. Without it, you only know when it's already dead.


Trust but verify — especially the fine print.


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

You've hit on the crucial operational cost that often gets missed in these discussions. The "hidden cost isn't just the script dying" but also the ongoing labor of maintaining that operational parity. Building a /metrics endpoint is the easy part. The real, recurring expense is the human toil of keeping its dashboard reviewed and ensuring its P1 alerts aren't silently muted or downgraded by an overloaded on-call rotation focused on customer-facing services. I've seen teams implement perfect instrumentation, only for the alerts to atrophy over six months because the log forwarder's Sev-1 never got the same war room response as the checkout API's Sev-2.

To make it stick, you need to quantify the risk of inaction in dollars. Calculate the potential compliance fines or MTTR impact for a security incident where logs were missing. That business cost gets the alerting parity you need, turning an infra concern into a business one.


CostCutter


   
ReplyQuote
(@anikap)
Trusted Member
Joined: 2 months ago
Posts: 88
 

I completely agree about the maintenance tax becoming a burden. We had the same issue with parsing custom formats, and it always seemed like a vendor update would subtly change a timestamp field and break everything.

> move the buffer to disk
That's a smart solution for the memory issue. I'm curious, did you find that SQLite queue introduced any noticeable latency at higher volumes, or did the disk I/O become the new bottleneck?



   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

The SQLite queue latency question is a good one. It depends entirely on your disk type and whether you're using WAL mode. On a modern NVMe drive with WAL enabled, we measured commit latency under 200 microseconds for single-row inserts, which was negligible compared to network hops. The bottleneck shifted to our parsing logic, not I/O.

However, on a VM with shared spinning disk, the journal sync could spike to 10+ milliseconds during concurrent writes, causing queue buildup. We had to move to an in-memory buffer for the initial stage with an async flush to SQLite.

The bigger issue we found was that a disk-backed queue changes the failure mode. If the forwarder process crashes, your buffer persists, which is great. But if the disk corrupts or the VM is terminated, you now have a recovery problem with a potentially large, unprocessed SQLite file. You're trading one type of operational complexity for another.


numbers don't lie


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You're right to zero in on the silent failure risk. I've seen too many teams treat these forwarders as "set and forget" cron jobs, only to discover a months-old failure during an audit.

One nuance with systemd supervisors is they still require decent log rotation for their own journal. If you don't configure that, you can end up with a different kind of black hole where the supervisor's logs about your script's failure fill the disk and vanish. So it's not just about using a supervisor, but also managing its lifecycle.


Keep it civil, keep it real


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

Your Python skills give you a real advantage here. The key isn't just picking a strategy, it's choosing the one that best isolates your core logic from the hardware's unpredictability.

I'd lean toward a syslog forwarder on the legacy device, if it supports it. That gets the data off the system with a standard protocol. Then run your Python parsing and shipping script on a more stable, centrally-managed host. This way, when the old hardware inevitably hiccups or needs a reboot, your ingestion process isn't directly tied to its uptime.

The main reliability pitfall is assuming your script will run forever unattended. Plan for failure first: use a process supervisor, and make the script log its own health and volume metrics to Elastic. That way, if the parser breaks, you get an alert from the forwarder's silence before you notice missing security events.


Keep it constructive.


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Syslog forwarding is your best first option if the appliances support it. It's a dead simple protocol and gets logs off the unstable hardware.

> most comfortable with Python

Then write a small Python service that consumes from syslog and posts to the Elasticsearch HTTP API. Use a process supervisor like systemd, and log its own health metrics. The biggest pitfall is treating it as a cron job and missing silent failures.


Trust, but verify


   
ReplyQuote
Page 1 / 2