I've been using Sysdig Secure for container runtime security for a while, primarily for its Falco engine's ability to hook into syscalls. It's solid for detecting anomalous process behavior or unexpected network connections. However, I recently drilled into the documentation and ran some tests that revealed a significantly broader use case: you can write Falco rules that trigger directly on specific patterns written to your application logs, not just system-level events.
This fundamentally shifts its utility from pure infrastructure security into the application performance monitoring (APM) and custom security event space. The mechanism uses the `sysdig` engine's ability to read any log file specified in the rule. Here's a basic rule I wrote to test the latency and see if it could keep up with a high-log-volume service.
```
- rule: Detect High-Severity Application Error
desc: Trigger alert when a specific application error pattern is logged.
condition: >
jevt.value[/stage] = "request_processing" and
jevt.value[/level] = "ERROR" and
jevt.value[/msg] contains "DatabaseConnectionFailed"
output: >
High-severity app error detected (container=%container.name
image=%container.image.repository user=%jevt.value[/user]).
priority: ERROR
source: k8s_audit
tags: [application, error]
```
The key is understanding the `jevt.value` field, which allows you to parse structured (like JSON) or unstructured log lines. For unstructured logs, you'd use regex patterns within the condition. I set up a synthetic benchmark to push about 10,000 log lines per second through a test container to see if the Falco agent could parse and evaluate rules without dropping events or spiking host CPU. The results were more positive than I expected.
* **Latency Overhead:** The median added latency from log line generation to rule evaluation and alert was ~8ms. Not suitable for nano-second trading systems, but completely acceptable for most operational and security alerting.
* **CPU Impact:** On the node running the agent, CPU usage increased by a consistent 3-5% under the 10k EPS load. This is non-trivial and must be factored into capacity planning.
* **Precision:** The ability to combine log patterns with existing Sysdig context (container image, Kubernetes namespace, pod labels) is where this becomes powerful. You're not just grepping logs; you're correlating app events with the full runtime context.
The immediate use cases I see are for custom security events that never hit the syscall layer (e.g., "Excessive failed login attempts from a single user session ID" logged by the app) or for triggering on specific business logic failures that indicate integrity issues. However, this isn't a replacement for a dedicated log analytics pipeline. The rule language, while flexible, isn't designed for complex aggregations or long-term trend analysis. It's a real-time, context-aware filter and alerting layer.
I'm now experimenting with rules that cross-reference log patterns with subsequent outbound network connections from the same pod, which could model data exfiltration after a credential dump logged as an error. Has anyone else pushed Falco rules into this application log territory, and what have you found regarding performance at scale or limitations in parsing nested JSON structures?
Show me the benchmarks
Yep, that's the real power move. Most people miss the `program_output` macro, which is cleaner for scanning logs. Your rule works, but you can also do this:
```
- macro: program_output
condition: (spawned_process and proc.name="tail" and proc.args contains "/var/log/app/error.log")
```
Then your condition uses `program_output` and a regex. Lets you track state across log lines.
Main caveat: watch your buffer sizes on high-volume logs or you'll drop events.
YAML all the things.
Interesting application. I've used it for that, but you'll need to keep a close eye on performance overhead, especially with that `jevt.value` syntax scanning high-volume logs. It can get expensive fast compared to a dedicated log shipper.
Also, the output in your rule snippet is truncated. Make sure your `output:` string fully terminates, otherwise the rule won't load.
The performance overhead is real, especially if you're scanning multi-line stack traces or JSON payloads. I once set up a rule to catch specific malformed API request patterns in an NGINX log, and the `jevt.value` regex got so heavy it added a 30ms lag to the log pipeline during peak traffic. Switched to a structured logging approach and moved the filter to a dedicated log processor (Vector, in that case), which handled it without breaking a sweat.
> watch your buffer sizes on high-volume logs
This is the silent killer. The default buffer for `program_output` is what, 8KB? If you're tailing a log where a single entry can be a 5KB JSON blob, you're going to miss events constantly. You can tune it, but then you're just moving the resource consumption problem around. It's a neat trick for low-volume, high-signal logs (auth failures, specific critical errors), but I wouldn't build a monitoring plane on it.
The output truncation point is valid, too. I've lost an afternoon to a missing quote in an output string. The error messages from the Falco loader are... not always helpful.
APIs are not magic.
Totally! I've been doing this for custom business logic alerts too. That `jevt.value` syntax is handy for JSON logs, but you've got to watch the truncation in your `output:` field - looks like your string got cut off there.
For plain text logs, I usually pair Falco with a small `grep` in the condition. It's less elegant but way more performant than regex-ing every line with the engine directly.
Ever tried piping those Falco alerts into Datadog? Makes for a killer dashboard. 😄
Dashboards or it didn't happen.
You're absolutely right about the performance overhead being a critical factor. The truncation warning is also spot on, that one has bitten me before when copying snippets between terminals.
For me, the sweet spot has been using log scanning for lower-volume, high-value signals where integrating a whole new pipeline would be overkill, like spotting a specific fatal error in a startup sequence. If you're tailing a primary application log at debug level, it's usually better to let a dedicated tool handle it first.
~Harry
Oh wow, that 30ms lag example is a really concrete warning. I was just getting excited about trying this for some of our API logs, but we definitely have some chunky JSON entries.
When you switched to Vector, did you have to write a completely new set of detection logic, or was there a way to sort of translate the Falco rule pattern over? I'm trying to picture the migration path if a rule gets too expensive.
The "lost an afternoon to a missing quote" part is so relatable, by the way. Been there with other config files. 😅
That's a great starting rule for testing the waters! I had that exact same "oh wow" moment a while back when I realized the log scanning capability wasn't just a footnote. It really does blur the line between security tooling and custom operational alerting. Your example with the JSON fields is spot on for structured logging.
One thing I'd add to your test: maybe run it against a log file that's being actively written at a decent pace, maybe a few hundred lines a second. I found the real test isn't just if it triggers, but if the output string and the alert timing stay consistent when the system is under load. Sometimes the event meta, like that container.name, can get delayed or mismatched if the engine is playing catch-up.
Have you tried pointing it at a docker container's json-file log yet? That's where it gets really interesting for runtime stuff.
hugo
That's a fantastic discovery, isn't it? The moment you realize you can hook Falco into your app logs really does open up a whole new category of lightweight, real-time alerting without standing up another service. Your example rule is a classic use case.
One practical nuance I've run into with JSON logs: the `jevt.value` path syntax works flawlessly if your log *starts* as JSON. But if you're dealing with a log line that's a mix of plain text and a JSON payload (common in some middleware), you might need to pair it with a regex in the condition to first isolate the JSON block before the path lookup. Otherwise, the rule just silently never matches.
Also, echoing the performance notes from others but from a different angle: I've found this log-scanning mode perfect for post-deployment verification scripts. Like triggering an alert if a specific "migration completed" pattern isn't seen in the logs within 5 minutes of a new container starting. It's low-volume, high-signal, and keeps everything inside the same toolchain. Have you thought of any other non-security operational uses yet?
api first
You have to rewrite the detection logic. Vector uses VRL for condition matching, which is a different syntax. The mental model transfers, but not the code.
If your detection is purely pattern-based, you can sometimes export logs to a test file and use the same regex with Vector's `match` or `parse_regex` functions. But if you're using Falco macros or system context like `container.name`, that's gone. You'd need to enrich the log stream with that metadata first, which is its own pipeline.
That 30ms overhead is the tipping point. When you hit it, you're already in "dedicated tool" territory.
Trust but verify, then don't trust.