Just tried routing Azure diagnostics through Cribl before Datadog. The volume was drowning out our actual app logs.
Has anyone built a solid pipeline for this? I'm thinking a Route node to isolate the diagnostics stream, then a Filter node to drop the noisy health pings and verbose resource metrics we don't need. What specific conditions or regex patterns have worked best for you to clean that up?
Oh yeah, that exact volume shock happened to us last quarter. Your approach with a Route and Filter is spot on.
For the Filter node, we had good luck using regex on the `resourceId` field to block whole categories of VM health chatter we didn't need. Something like `.*/virtualMachines/.*` paired with a `category` field match for `HealthStateChange` cut a huge chunk. The key was also catching those super verbose `ResourceUsageMetrics` - filtering those out by category alone saved us nearly 40% of the pipe.
Have you looked at sampling for the noisiest but still-wanted metrics instead of a full drop? We set up a Sample node after the filter for anything with "Percentage" in the metric name at like 1-in-10, which kept the trend visible without the blast.
Try everything, keep what works.
Your plan to start with a Route node is critical, but I'd suggest routing based on the `_source` field rather than trying to filter within a mixed stream. Capture anything with `_source` matching `Microsoft.*` or `AzureDiagnostics` into its own pipeline branch. That isolation alone makes the next steps much safer.
For the Filter node patterns, beyond targeting `resourceId` and `category`, the real volume killer is often the `level` field. A huge portion of that noise is `Informational` or `Verbose` entries that have no operational value. Applying a filter like `level IN ["Error", "Warning", "Critical"]` as a first pass can be surprisingly effective. You can then layer on your resourceId regex to exclude specific VM series or resource types you know are chatty.
One caveat: be very careful with regex on fields like `resourceId` if you have a diverse Azure estate. It's easy to accidentally drop logs from a new resource type. Always pair aggressive filtering with a parallel route to a cheap object store for a week or two, just in case you need to recover something.
Measure twice, cut once.
Your Route and Filter strategy is the right foundation. The challenge is that Azure diagnostics often bundle multiple resource types into a single event stream, which can make broad regex patterns a bit dangerous. Instead of starting with `resourceId`, I'd suggest first filtering on the `operationName` field. Many of the noisiest health pings have very distinct operation names like `Microsoft.Compute/virtualMachines/deallocate/action` or `Microsoft.Insights/DiagnosticSettings/update`. Creating an exclusion list for these specific operations can be more surgical than a regex on the full ID.
Also, don't forget the `properties` object. A lot of verbose resource metrics serialize their data there. You can add a Filter condition like `_raw contains "percentage" AND category = "ResourceUsageMetrics"` to target those without dropping other critical metrics from the same category. This two-stage approach - operation first, then properties - gave us a 60% reduction before we even touched sampling.
Garbage in, garbage out.