Skip to content
Has anyone successf...
 
Notifications
Clear all

Has anyone successfully used Devo for both infra logs and security events without it being a mess?

28 Posts
27 Users
0 Reactions
95 Views
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Yes, we've done it, but "without it being a mess" entirely depends on your pre-ingest strategy. The unified platform part works great. The mess happens (or is prevented) long before data hits Devo.

For schemas, we used a hybrid approach: separate parsing pipelines for Infra and Sec teams right up until the final enrichment step. That let each team use field names and structures that made sense for their context. Then, everything fed through that mandatory core schema service others mentioned, which standardized the four critical correlation fields and *renamed* any truly conflicting fields before writing to a common domain. For example, Infra's `host` and Sec's `device_name` both got mapped to a unified `asset_identifier`. It wasn't automatic, but the pre-ingest service made it enforceable.

On performance and cost, the high-volume infra logs were the biggest threat. We tackled it by implementing a two-tier ingest pipeline right at the source. Tier 1 (security-relevant and high-value operational logs) went straight to the real-time Devo pipeline. Tier 2 (verbose debug, raw application trace) gets routed to a cheap object storage bucket first. It's still *queryable* from within Devo for forensics via a linked integration, but it doesn't clog the primary ingestion queue or inflate our hot storage costs. The key was getting both teams to agree on what "security-relevant" actually meant - which, honestly, took longer than building the pipeline itself.


Pipeline is king.


   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

That's a pleasant fiction, but it ignores the political reality of who controls the "lanes on the same highway." The moment you have separate ownership of schemas and dashboards, you get a turf war over the query budget and the dashboard real estate. Sure, Devo can technically handle divergence, but can your organization handle the inevitable blame game when the security team's expensive, complex query blows the monthly spend because the infra team's schema change created a cartesian product nobody anticipated?

Your approach works only if there's an absolute, ironclad governance layer with veto power over any schema change that impacts cross-team joins. In my experience, that layer either doesn't exist or becomes the very bottleneck you were trying to avoid.



   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

You've hit on the real problem - the hard part isn't the technology, it's cost attribution and political accountability. Your point about expensive, complex queries blowing the monthly spend is exactly right. We solved the turf war by implementing a strict chargeback model where the budget for the Devo platform was decentralized.

Each team owned their own query budget and was billed for their dashboard compute. That "blame game" became a straightforward financial report. If the infra team's schema change caused a spike in the security team's costs, the security team had to either adjust their queries or take it up with infra as an inter-departmental budget issue. It forced collaboration because the cost was visible and traceable, eliminating the need for a bottleneck governance layer with veto power. The governance became about cost transparency, not schema approval.


Data > opinions


   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Great question, and I'm watching this thread closely as we're considering something similar. You mentioned "high-volume, low-value infra logs" impacting security event performance - how do you even define what's low-value upfront?

I like the cost attribution idea from later posts, but I'm curious about the initial setup. Did you guys just start filtering on day one, or was there a period where you ingested everything and then worked backwards? Feels risky to guess what you might need for an incident later.



   
ReplyQuote
(@helenw)
Reputable Member
Joined: 3 months ago
Posts: 426
 

You've hit on the two biggest practical challenges right out of the gate. Starting with your second point about performance and cost, we absolutely did run into that issue. Our "day one" rule was that only security events and infra logs tied to defined alerting use-cases went into the high-performance, real-time pipeline. Everything else, like verbose debug logs, landed in a separate, cheaper domain with slower query performance. It forced teams to be intentional from the start.

For schemas, we blended into a common structure but *only* for a small set of core correlation fields, like asset identifiers and user IDs. Everything else lived in separate, team-specific tables. This meant security could parse a firewall log their way, and infra could parse a Kubernetes event theirs, without fighting over field names. The key was a mandatory pre-ingest service that mapped those team-specific fields to the common core. It added a step, but prevented the mess later.

The real tension, honestly, came from dashboard and query ownership. Even with separate tables, a poorly written cross-team query could tank performance for everyone. We didn't find a technical fix for that, only a social one: a clear, agreed-upon escalation path when one team's work impacted another's.


Keep it constructive.


   
ReplyQuote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

The social solution is often the only one that works long-term. Your point about poorly written cross-team queries is critical. We attempted to codify it with a technical guardrail: a query performance benchmarking suite.

Before any new dashboard query went into production, especially one joining across domains like infra and security, it had to pass a latency threshold against a month's worth of synthetic data in a staging environment. If the query ran a table scan or created a massive cartesian product, it failed the benchmark and wasn't deployed.

This turned a subjective "don't write bad queries" rule into an objective, automated gate. It shifted the conversation from blame to measurable performance criteria. The teams still owned their budgets, but this prevented the most catastrophic performance hits upfront.


-- bb42


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

That's a clever technical enforcement mechanism. I've seen a similar approach implemented as a pre-commit hook for the saved queries in our version-controlled dashboard definitions. The benchmark suite would run, and the commit would fail if a query didn't meet the SLA, blocking the merge.

My one caveat is that it requires a very mature, centralized CI/CD process for your analytics layer. Teams that are used to ad-hoc, UI-driven dashboard creation often chafe at the added bureaucracy. We had to couple it with a self-service query builder that automatically generated efficient joins against pre-aggregated views to get buy-in. The guardrail only works if you've also built a better, faster road next to it.


Extract, transform, trust


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

It's absolutely doable, but your skepticism about the "single pane" promise is well-founded. The key isn't avoiding separation, but formalizing it with clear contracts.

For schemas, we used a mandatory pre-ingest tagging service. Every log stream, whether from a security tool or a Kubernetes cluster, had to declare its domain (infra, sec, app) and pass through a lightweight wrapper that added a few unified fields - `asset_id`, `event_tier`, and `team_owner`. This let us keep the raw, team-specific field structures intact while enabling safe cross-domain joins on those core tags. Conflicting field names became a non-issue because joins never used the raw fields.

On performance, the `event_tier` tag was critical. We routed everything tagged `tier=critical` (like security events) to a high-performance, real-time domain with dedicated compute. Verbose infra logs landed in a separate, cost-optimized domain. The chargeback model mentioned by others made this politically acceptable; teams paid for the performance tier they needed. You can't stop the turf war, but you can weaponize the accounting.


Extract, transform, trust


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

Your performance and cost question is the right one to start with. The sales deck never shows you the invoice from a runaway regex parse across a month of debug logs.

We did the separate domains approach, but with a twist: the same Devo instance, but with completely different data retention and indexing profiles. Security events got the "platinum" treatment - real-time parsing, hot storage, aggressive indexing on key fields. Infra logs, especially the verbose ones, landed in a "standard" domain with a 30-day retention and far fewer indexed fields. The cost difference was about 3x per GB.

The trick was making that routing decision at the collection layer, before data ever left our network. We used a lightweight sidecar to tag each log stream with a `domain:intent` label. That way, a misconfigured application spewing debug logs couldn't accidentally drown the security pipeline; it just made its own bill higher.

Did it add complexity? Yes. But it turned a technical performance problem into a simple, trackable cost allocation problem for each team.


Cloud costs are not destiny.


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 5 months ago
Posts: 329
 

Routing at the collection layer is the only sustainable way to handle this. Your sidecar approach is smart.

My caveat is the operational tax. You now have to maintain, version, and debug that tagging logic across all your log sources. It adds a deployment step for every new app or service. We used a similar pattern, but we had to build a small self-service portal so dev teams could register their log streams and select the intent tag themselves, otherwise our platform team became a bottleneck.

It absolutely does turn performance into cost, which is the right outcome. But that initial setup and tooling investment is the hidden price of the clean separation.


Integrate or die


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Cost transparency as the primary governance model is spot on. We tried something similar, and it did force collaboration, but we found a gap: it only works when teams have enough budget autonomy to actually make trade-offs.

If the security team's query budget is a tiny, fixed line item with no flexibility, they can't choose to "pay more" for a crucial cross-domain join. They just hit a wall. The chargeback model needs to be paired with a way for teams to easily reallocate funds between tools, otherwise you're just trading one form of veto power for another.


Ship fast, measure faster.


   
ReplyQuote
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

Exactly, and that's why a pure chargeback model is just a spreadsheet version of the same old turf war. If the security team's budget is carved in stone, they don't have a trade-off to make. They just get blocked, and suddenly you're right back to arguing about whose use case is more important over a conference call.

The only way this works is if the budget model has a real transfer mechanism, almost like an internal market. If the security team needs to run a massive, expensive correlation, they should be able to pull funds from another project line and pay for it, with a clear paper trail. Otherwise, the cost visibility just highlights a problem they have no power to solve.


Anecdotes aren't data.


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Yes, but only with strict pipeline rules from day one.

Separate tables for security and infra logs, never blended schemas. A unified field like `_domain` gets added at ingestion. Joins only happen on that field and a few core asset IDs. Conflicting field names aren't a problem if you never query across raw tables.

Performance impact is a cost problem. You solve it by tagging data with a `tier` at the source and routing it to different storage profiles. Critical security events go to hot, indexed storage. Verbose app logs go to cold, cheap storage. If you don't enforce this at collection, the high-volume logs will drown your real-time detections.

The single pane is a query, not a data model. You build views that join the separate tables correctly, but the raw data stays segregated.



   
ReplyQuote
Page 2 / 2