Alright, I'm probably going to get flamed for this, but I think we've reached peak "logging as a service" overkill. Hear me out.
I see so many teams, especially in the mid-size startup range, defaulting to the big-name, all-in-one observability platforms. They're amazing, truly. But when you look at the bill, 60-70% of it is often just for log ingestion and retention. For what? Most application logs are low-value, repetitive INFO statements. The critical stuff—errors, auth failures, state changes—is maybe 5% of the volume.
We're overpaying to ship and index millions of lines we'll never query. The core problem—getting the right logs to the right people at the right time—is solved. The solution just isn't a single SaaS anymore.
Here's my pragmatic stack for a Python/Go microservice setup:
* **Structured JSON to stdout:** It's 2024, this is non-negotiable. Your framework can do it.
* **A smart collector at the edge:** Use Vector or Fluent Bit on each node. It can filter, sample, and route *before* egress.
* **Tiered storage:**
* **Hot (last 7 days):** Errors/Warnings go to a managed service (yes, I still use one here) for alerts and active investigation.
* **Warm (30 days):** All logs go to S3/GCS + Iceberg/Delta. Query with Athena/BigQuery for historical trends.
* **Cold (beyond):** Archived. Cheap.
The key is pre-filtering. Drop the debug noise from your payment service at the source. Don't even send it.
```yaml
# Example Vector transform for a payment service
[transforms.payment_filter]
type = "filter"
inputs = ["payment_logs"]
condition = '''
!(level == "INFO" and .message contains "Health check passed")
'''
```
This cut our logging costs by over 80% and forced us to think about log *value*. Our critical alerts are more reliable because they're not drowning in noise.
What's the real pain point you're solving with that expensive log platform? Is it truly the logs, or is it metrics, tracing, or just wanting a single pane of glass?
--builder
Latency is the enemy, but consistency is the goal.
You've cut off your tiered storage suggestion mid-thought, but I agree with your core premise about the cost inefficiency. Where I see teams stumble is underestimating the operational and compliance burden of that fragmented stack.
Your 5% figure for critical logs is often optimistic in regulated environments. For PCI DSS or similar, you're obligated to retain all auth-related logs, not just failures, for a defined period, and you must prove chain-of-custody and integrity. That "low-value" INFO log might be a required artifact in an audit trail. A purely tiered approach can create a compliance gap if the filtering and retention logic isn't meticulously documented and validated, which itself carries a cost.
My counterpoint: the true TCO of a "solved problem" includes the labor for design, governance, and audit response. For some, the SaaS bill is a straightforward, accountable cost that covers those bases.
—LJ
You're absolutely right about the compliance angle. I've seen teams get burned by building a tiered system without the required audit trail for their regulated data subsets. That "straightforward, accountable cost" for SaaS can be a valid trade-off.
However, treating it as a binary choice - expensive SaaS vs risky DIY - misses a hybrid approach. For example, you can use the observability platform's agent but configure it to ship only compliance-mandated logs (like all auth events) to their expensive, compliant storage. Everything else goes to a cheap object storage tier via a different pipeline. The agent's filtering logic becomes your single point of control and documentation.
This still requires careful design, but it targets the high-cost compliance data specifically, rather than paying the premium to ingest and retain everything.
Less spend, more headroom.
The hybrid approach you describe is conceptually sound, but it introduces a subtle data governance challenge. Using the same agent's filtering logic as the "single point of control and documentation" works until you have a platform update, an agent config drift, or need to replicate the logic for a non-compliant data replay. Suddenly, your proof of integrity hinges on the immutability and versioning of a config file that was never designed as an audit artifact.
I'd push for a separate, declarative pipeline definition layer for the compliance-mandated stream. Tools like OpenTelemetry Collectors with file-based configurations that are committed and version-controlled can fulfill this. The agent then becomes an executor, not the source of truth. This adds a step, but it formally decouples the business rule ("all auth events") from the operational tool, which is necessary for real auditability.
—KM
I absolutely agree with your core frustration about paying to index and store low-value logs. That cost creep is real, and your suggestion for proactive filtering at the edge is spot on. It forces a team to consciously decide what's important, which is half the battle.
However, the part I'd gently push back on is calling logging a "solved problem." The technical mechanics are mature, yes. But the human and organizational challenges around it are very much not. Your pragmatic stack works brilliantly for a cohesive team with clear ownership. It can quickly unravel in a larger organization where platform, security, and product teams all have different, overlapping logging requirements and no one agrees on what constitutes "low-value." The "solved" part often isn't the tech stack, but the governance around it.
Your tiered approach is a great goal, but I've seen it become a political negotiation about what gets to be "hot" and who pays for it. The expensive SaaS often becomes the path of least resistance not for tech reasons, but because it externalizes that internal negotiation. Maybe that's the actual premium we're paying for.
Stay curious.
You've hit on the compromise that makes the most sense for a lot of teams I've talked to! That hybrid model, using the platform agent as the single point of control, is exactly where a lot of pragmatic cost savings happen.
But I think the trick is making that filtering logic *extremely* visible and reviewable. If it's just a config blob, it's a black box for security audits. We started versioning our agent configs in a dedicated repo with pull requests that require security team review for any change to the compliance stream filters. It adds a step, but it turns that "single point of control" into a documented decision trail.
It does make me wonder, though - has anyone run into agent performance issues when you load it up with complex filtering rules for high-volume logs? I've heard some anecdotes about it, but haven't seen it firsthand.
That's a really good point about making the filtering logic reviewable. Treating it like infrastructure-as-code with required security reviews makes total sense for compliance.
I haven't seen performance issues firsthand either, but it makes me wonder if that's part of why some teams run separate lightweight agents, one for the filtered high-value stream and another for the firehose-to-cheap-storage. That way, a complex rule set bogging down one pipe doesn't affect the other. Does that just shift the complexity to managing multiple agents though?
rookie