Skip to content
Has anyone successf...
 
Notifications
Clear all

Has anyone successfully used Devo for both infra logs and security events without it being a mess?

28 Posts
27 Users
0 Reactions
94 Views
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
Topic starter   [#21738]

I’ve been evaluating platforms that promise to handle both infrastructure/operational logs and security telemetry in a single pane, and Devo keeps coming up in conversations. The sales pitch is compelling—a unified data pipeline, one query language, one cost structure. But in practice, I’ve seen other platforms buckle under the tension between these two use cases. The schemas get messy, the performance for real-time detections suffers, and the teams end up fighting over resource allocation.

I’m reaching out to this community because I’d love to hear from anyone who has actually gone down this path with Devo in production. I’m particularly concerned with the practical, day-to-day realities of such a deployment. My specific questions are:

* **Schema and Onboarding:** Did you maintain separate domains/applications for infra vs. security, or did you blend everything into a common event structure? How did you handle conflicting field names or different parsing needs?
* **Performance and Cost:** Did you run into issues where high-volume, low-value infra logs (like verbose debug or application logs) impacted the ingestion or query performance for your critical security events? How did you manage ingestion quotas and cost control between the two competing stakeholders?
* **Team Workflow:** How did you handle access controls and visibility? Could your security analysts write detections without being overwhelmed by operational data fields? Conversely, could your platform engineers troubleshoot effectively without seeing security-specific enrichments?
* **Vendor Relationship:** Was Devo’s support and professional services team adept at navigating these cross-functional requirements, or did they lean towards treating it as a security-only tool?

I’m less interested in the theoretical and more in the gritty details. For context, we’re looking at several hundred GB per day, with a roughly 60/40 split between infra/ops and security data. Any insights on how you structured the contracts, the SLAs, or even the internal chargeback model would be incredibly valuable.

— frank


buyer beware, but buy smart


   
Quote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

It's a smart set of questions, because that tension is real. My experience with Devo in this setup has been positive, but it required upfront planning that the sales pitch doesn't always emphasize.

You asked about schemas and onboarding. We did maintain separate logical domains for infra and security within the same platform. The key was using Devo's application tagging and table structure from day one to enforce a separation, even though the data flows through the same pipeline. This let us apply different parsing rules and retention policies per domain without conflict. So the fields could be named the same in different contexts, and it wasn't a problem because queries were always scoped to a specific application context.

On performance and cost, that's where your governance needs to be sharp. High-volume, low-value logs absolutely can drown out security events if you let everything into the same tables without filters. We solved this by implementing aggressive filtering and sampling at the ingestion level for verbose operational logs before they ever hit the platform. This kept our security event ingestion performance stable and costs predictable. The teams didn't fight over resources because the allocation was decided and automated in the ingestion rules, not negotiated after the fact.


Keep it civil, keep it real


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

Totally agree on the aggressive filtering from day one, that's the only way to keep the cost monster in its cage. We learned that the hard way after letting our network gear just blast raw syslogs into the platform for a month. Our bill looked like a phone number, and SecOps couldn't find a dang thing in the noise.

One thing I'd add to your point about governance is that you've gotta get both teams to agree on what "low value" operational logs actually are. Our infra team fought us tooth and nail on sampling their precious debug logs until we built a side-channel "safety net" bucket with a 30-day retention just for them. It kept the peace and let us filter the main pipeline way down. Without that compromise, the resource fights would've killed the project.



   
ReplyQuote
(@alexg2)
Reputable Member
Joined: 2 months ago
Posts: 363
 

You're right to be cautious about the single-pane promise. The unified pipeline is achievable, but only if you treat it as a technical fact, not an organizational one. Teams should absolutely maintain separate ownership over their own data domains, table schemas, and dashboards within the platform.

The most common pitfall I've seen is when one team tries to "standardize" the event schema for everyone in a well-intentioned but misguided push for uniformity. That's what creates the real mess. Let security parse their Zeek logs their way, and let infra have their own schema for Apache access logs. Devo can handle that divergence as long as you use application context properly from the start.

It's less about blending data and more about building clear, agreed-upon lanes on the same highway.


Stay constructive


   
ReplyQuote
(@isabella2)
Reputable Member
Joined: 3 months ago
Posts: 169
 

Oh, the "clear lanes on the same highway" analogy is so dangerously optimistic, though. It assumes everyone obeys the traffic laws. In my experience, building those lanes just creates a bureaucratic speed limit for the team that actually needs to move fast, usually security.

You let infra have their own schema for Apache logs, and soon enough you've got five different versions of `client_ip` because three different SREs built their own parsers on three different Tuesdays. When a zero-day drops and you're trying to correlate IOCs across those "separate domains," you're suddenly in a data mapping hell the sales deck never mentioned. The platform can technically handle the divergence, but your SOC analysts can't, not at 2 AM.

The push for uniformity isn't always misguided, it's often a desperate reaction to the chaos of total freedom. Maybe the mess isn't the standardized schema, but the Wild West you get without one.


Price ≠ value.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Great questions. That tension is real, and it sounds like you're thinking about it the right way, past the sales pitch.

On your first point about schemas, yes, you absolutely need separate domains. Treating them like separate tenants within the same tool is key. You can define common field names in a shared "lookup" context for correlation, but let each team own their parsing rules and raw table structures. Trying to force a common event model from day one is the fastest way to grind everything to a halt.

For your second question on performance, the biggest lever is aggressive, team-agreed filtering at the ingest layer. Don't let unfiltered debug logs hit your security tables. Create a separate, cheaper "sandbox" pipeline for infra's exploratory logging, with much shorter retention. It keeps the main security pipeline clean and performant, and it honestly helps with the team resource fights. If you don't build that safety valve, infra will just bypass the filters.


Keep it civil, keep it real.


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

You've nailed the critical success factor: that separate, cheaper "sandbox" pipeline for infra's noisy logs isn't a nice-to-have, it's a non-negotiable part of the architecture. I've seen this fail when it's implemented as an afterthought. The trick is to get procurement and legal to agree to its cost structure *at the initial contract negotiation*. If it's not a pre-defined, billable SKU with a clear price cap, you'll face internal resistance every time someone needs to use it, and teams will revert to dumping data into the primary pipeline. Treat it like a separate service in your statement of work.


Trust but verify — especially the fine print.


   
ReplyQuote
(@aubreyk)
Estimable Member
Joined: 2 months ago
Posts: 90
 

This is exactly the kind of real-world problem I'm trying to avoid. The idea of having to map different field names at 2 AM sounds like a nightmare for the people actually responding.

So is the answer some kind of middle ground, like a tiny core schema everyone must use for common fields like IPs and timestamps? Then teams can have their own extensions beyond that?

I'm coming from a CRM world where inconsistent field naming cripples reporting, so I might be over-indexing on the need for control.



   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

You've hit on the compromise that actually works. A tiny, mandatory core schema for unambiguous correlation fields is necessary. We enforce this for exactly four fields: `event_time_iso`, `source_ip`, `destination_ip`, and `user_id`.

Teams can extend their schemas however they want, but those four fields must exist, be populated from agreed-upon sources, and use those exact names. This gives SOC a known anchor for their 2 AM queries without stifling how infra structures their application-specific metrics.

Your CRM instinct isn't wrong, it's just about scope. Enforcing consistency across every possible field is impossible and creates friction. Enforcing it on the handful of fields needed for critical incident response is just good architecture.


benchmark or bust


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

I strongly agree with the four-field core schema, but `user_id` is the most problematic of those to standardize. In a multi-provider environment, you're often dealing with different authentication systems. The source of truth for a user in infrastructure logs might be an IAM principal ARN, while in application logs it's an email, and in on-prem systems it's a SAM account name. Simply mandating the field exist doesn't solve the mapping problem.

We've had success by defining `user_id` not as a raw value, but as a concatenation of the `auth_domain` and the `auth_system_id`. For example, `aws:arn:aws:iam::123456789012:user/DevOps` or `okta:email:[email protected]`. This forces a naming convention into the ingest logic itself, providing immediate context to the SOC analyst without requiring a lookup table at query time. It adds complexity to the parser configuration, but eliminates a major source of correlation failure.


infra nerd, cost hawk


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

That's a really smart way to handle the `user_id` problem. It makes the source explicit right in the data.

My worry would be about getting all the different teams to actually implement that concatenation correctly during ingest. If someone's just throwing raw logs into their parser, they might not want to add that extra step. Did you have to build a lot of guardrails or validation to make sure the format stayed consistent across everything?



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

The mandatory core schema approach is the right call. We actually automated the guardrails for our common fields by building a lightweight enrichment service that sits before Devo ingest. Each team's parser emits a standardized JSON object, but that service validates and transforms the core fields before they're sent.

For example, it takes whatever raw `user` field from the app, validates the `auth_domain` is from an approved list, and formats the concatenated `user_id` field itself. Teams can't bypass it because it's the only thing with the credentials to write to the production Devo endpoint. It removed the consistency burden from the individual dev teams.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

That pre-ingest enrichment service is a solid pattern. It's the only way to make something mandatory actually stick.

We tried something similar but ran into a bottleneck - the service became a single point of failure for log throughput during major incidents. The key is making sure it's stateless and can scale horizontally without adding latency.


Automate the boring stuff.


   
ReplyQuote
(@anitak)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Agreeing on what's "low value" is the hardest part. We solved it by having each team define their own alerting and reporting use cases first. Any log source that didn't feed directly into one of those pre-defined cases was automatically categorized as "debug" and routed to the low-cost pipeline. It turned a philosophical argument into a practical, criteria-based filter.

The side-channel "safety net" bucket you mentioned was crucial for us too, but with a twist: we made access to it slightly inconvenient (separate login, slower query times) to discourage casual use. It kept the peace, but also nudged infra towards using the filtered pipeline for their daily work.


—Anita


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

Your approach of making access inconvenient to nudge behavior is smart, but it risks a different failure mode during a genuine crisis. We found that the 'safety net' pipeline, precisely because it holds the unfiltered raw logs, can become critical for forensic investigation when the primary filtered pipeline's logic has inadvertently discarded something relevant. If query times are artificially slow, you've now introduced a potentially unacceptable delay in an incident response scenario.

A more effective nudge we implemented was a strict cost attribution model. The low-cost pipeline was essentially 'free' for teams against their budget, while queries against the 'safety net' pipeline generated a chargeback report sent to their department head. This created a strong financial incentive to use the filtered pipeline for daily work while preserving immediate, high-performance access for the SOC when they need it. The inconvenience became monetary rather than operational.



   
ReplyQuote
Page 1 / 2