We're looking at a migration where about 30% of our log sources are still on-prem, running on old hardware and some proprietary software. The cloud-native apps are straightforward with Sumo, but these legacy systems are a different story.
I've been testing a few methods for getting those logs into Sumo Logic without a major re-architecting project. Here's what's been somewhat practical so far:
* **Installing the collector on a central log forwarder VM:** This is the most common approach for us. We have a few Linux VMs acting as central log aggregation points. We install the Sumo collector there and use it to tail flat files shipped from the legacy systems via syslog or even scheduled SFTP scripts.
* The main challenge is parsing those non-standard formats. We've had to write a fair number of custom source categories and use Sumo's field extraction rules to make sense of the data.
* **Using syslog-ng/rsyslog as a relay:** For systems where we can't install anything, we configure them to send syslog to a relay server that has the Sumo collector installed. The collector then uses a Syslog source.
* This works, but you have to be careful about timestamp parsing and time zones from these old systems—they're not always consistent.
Here's a basic example of a `sources.json` config for the collector that tails an aggregated application log directory:
```json
{
"api.version": "v1",
"sources": [
{
"name": "Legacy-App-Logs",
"pathExpression": "/opt/legacy/logs/app_*.log",
"sourceType": "LocalFile",
"automaticDateParsing": true,
"multilineProcessingEnabled": true,
"useAutolineMatching": true,
"forceTimeZone": "America/New_York"
}
]
}
```
The real time sink isn't the collection, but structuring the data afterward. Has anyone else dealt with a similar hybrid setup? I'm particularly interested in:
* How you handle log rotation on the legacy systems to ensure no data loss.
* Any clever ways to enrich or tag this on-prem data differently in Sumo to keep cost attribution clear.
* If you've found a more agentless method that's reliable for truly locked-down systems.
Our bill shows a clear cost per GB for this data, so making sure we're only sending what we need is a priority.
terraform and chill
Your point about timestamp parsing with the syslog relay is crucial. In our audits, we often see events where the relay server's local time gets appended, creating mismatches with the original log event time. This breaks sequential analysis and can be a compliance headache for timelines.
We ended up mandating that all legacy systems configured for syslog must send their events in a format that includes the originating system's timestamp, even if it's a proprietary string. The parsing burden then moves to Sumo, but at least the source data is intact.
Have you considered the data residency implications if your relay VM is in a different region than the legacy systems? That was another surprise for us during a GDPR review.