Everyone talks about cutting ingest costs but most guides are fluff. Here's what actually works: drop the noise before it hits the collector.
Start by identifying your top 10 log sources by volume. For 80% of shops, it's load balancer logs, verbose debug output from apps, and repetitive health checks. Sumo's built-in metadata fields like `_sourceCategory` are your friend. Create an ingest budget rule that excludes anything with `_sourceCategory="/aws/elb/access"` and `_sourceHost` containing "healthcheck". That alone can trim 5-10%.
Next, use sampling for high-volume, low-value logs. Don't sample security events. Do sample verbose application debug logs at 1-in-10. Use a conditional statement in your collector configuration. If your log matches a high-noise pattern, apply the sampling rate.
Finally, kill the legacy syslog feeds from decommissioned systems. They're still running somewhere. Find them with a volume report by source IP and cut them off at the source. No point filtering what shouldn't be sent.
This isn't magic. It's basic hygiene. Do these three things and you'll hit 20% reduction without losing anything important. If you're not doing this, you're paying for garbage.
show me the logs
Your approach is fundamentally sound, especially the focus on top-volume sources and the pruning of legacy feeds. However, I'd caution against blanket exclusions like `_sourceCategory="/aws/elb/access"`. While it's a high-volume culprit, it often contains the only record of critical user-facing errors for web applications. A more nuanced method is to filter within that stream itself, perhaps excluding logs where the HTTP response code is 200 or where the request path matches a static health endpoint pattern. This preserves investigative capability while still cutting a significant portion.
Also, your 20% reduction target is reasonable, but I've found the actual attainable figure is highly dependent on organizational maturity. Teams with established FinOps practices often see diminishing returns after 15% from pure filtering, as the remaining "noise" is usually tied to legitimate, if verbose, application behavior. The next 5% typically requires engineering work to reduce log verbosity at the source, which has a higher TCO.
On sampling: conditional sampling in the collector config is powerful, but be meticulous with your pattern matching. A poorly scoped regex can inadvertently sample critical transactional logs. It's worth validating the sampled output in a staging environment for a full day before deploying to production.
Trust but verify.
Agreed on the blanket exclusion being risky. Your suggestion to filter within the stream using response codes or path patterns is spot on. I've done exactly that by creating a collector rule that drops ELB entries with `http_status=200` AND `request_path` containing `/api/health`. It cut about 70% of that category's volume but kept all the 4xx/5xx errors.
Your point about diminishing returns after 15% rings true. Once you've filtered the obvious noise, you're often left debating the business value of verbose debug logs from legacy services. Getting that last 5% reduction usually means convincing an app team to change their logging level, which is a whole different project.
Your point about killing legacy feeds is critical. I've seen teams run volume reports and find entire subnets still sending verbose logs from services that were sunset six months prior. It's pure waste.
A practical step after you've identified those feeds: don't just shut them off. Redirect them to a local syslog server for a two-week observation period first. Sometimes a forgotten monitoring script or cron job depends on that stream. This prevents a 3 a.m. page about a "critical service outage" that's just a dead log feed.
Your three-step method works, but the order matters. I always start with the legacy feed removal, as it's pure savings with zero analytical downside. That often gets you halfway to your 20% before you even touch filtering logic.