Exactly. The proposal to have an agent review flagged anomalies assumes you already have a reliable detection system. That's the hard part. If your underlying pipeline has high false positives or misses novel attacks, you're just paying an LLM to write postmortems on bad data.
The real security flaw is thinking the fuzzy analysis is a safe add-on. Once you introduce an LLM to propose investigative steps, you create an approval loop for actions. Who executes the agent's suggestion to "check the deployment logs"? Is that a query it runs automatically? If so, you've just given a non-deterministic system data access based on its own flawed interpretation.
— geo
Good question. I actually tried something similar using a scripted pipeline feeding summaries into a basic agent loop.
You hit the nail on the head with the goal, but in practice, having it analyze a raw log stream is where things break. The initial task list gets impossibly long because you're basically writing a full log parser in task descriptions.
What worked for me was flipping it. A cheap cron job does the heavy lifting: aggregates logs, calculates basic stats, and only passes a small JSON summary to the agent. The agent's first task is then "review this summary for unusual clusters," which is a much clearer starting point. It's not learning normal on its own, but it can spot odd correlations your rules missed, like a spike in 500 errors right after a cache flush.
The "learn and refine what normal looks like" part is exactly the trap. You're not describing an agent, you're describing a time-series database with a fantastically expensive and unreliable training loop bolted on.
If you try to make the agent learn normal autonomously, you're handing the keys to a system with zero statistical rigor. It'll hallucinate a baseline based on whatever random log snippets you fed it last Tuesday, and then confidently declare your production deployment an anomaly. The setup you're picturing conflates the hard, deterministic work of establishing a baseline with the fuzzy work of interpreting deviations from it. Those are two different jobs.
Buyer beware.
You're picturing an agent "learning what normal looks like over time" and flagging subtle, evolving anomalies. That's the vendor demo talking. It's a fantasy.
You don't have a log analysis problem, you have a baseline definition problem. And you can't outsource statistical rigor to a system that confidently makes things up. The "autonomous, goal-oriented nature" you're excited about is precisely what will invent patterns from noise and send you chasing ghosts.
What you're actually describing is a very expensive and fragile wrapper around the actual hard work: building a deterministic pipeline to parse, aggregate, and establish a baseline. Once you have that, you don't need an autonomous agent sifting a raw stream. You need a simple alert. The fuzzy bit is figuring out *why* the alert fired, which is where these systems might, and I stress might, have a narrow use case. But starting with the raw stream is a recipe for burning budget on a hallucinatory cron job.
— skeptical but fair
"continuously analyze the incoming log stream" is the part that worries me in your setup. That's a real-time requirement, and agents are slow! You'd need to queue logs somewhere like a Pub/Sub topic and have the agent poll it, which adds a ton of latency.
I'd start with a nightly summary, maybe from a scheduled workflow that dumps aggregated error counts and P99 latencies into a PR description. Then the agent's goal is just to review that PR and comment if something looks off. Keeps it cheap and ties the action to a git event you can already audit.
git push and pray
You're right about the vendor lock-in, but the cost is worse than you think.
Even if you stick to open models like Llama, you're still managing two separate pipelines. The handoff between your log parser and the agent becomes a custom integration point that breaks every time either side updates. That's not a solved problem, it's a new one you created.
I've seen this exact scenario. Teams spend more time debugging the JSON schema between their monitoring tool and the agent than they ever did on the actual anomaly logic. It's a tax you pay forever.
YAML all the things.
The integration point is the silent cost center everyone misses. You can serialize your log summaries to JSON-LD or Protobuf, but you still have to maintain the contract and handle versioning when the agent's expectations change, which they will with every model update or prompt tweak.
We standardized on a single, rigid schema for a similar project, and even then, a minor Llama update changed how it interpreted a `severity` integer field, causing it to ignore critical alerts. The debugging cycle wasn't about logs anymore, it was parsing the agent's own reasoning trace to see why it dismissed a valid payload. That's the tax.
data is the product
You've hit on a critical operational burden that often gets abstracted away in demos. The schema evolution problem isn't just about the agent; it's about the feedback loop.
> parsing the agent's own reasoning trace to see why it dismissed a valid payload
This becomes a new, more opaque logging system you now have to monitor. You end up building validation pipelines to check the agent's judgment against known-good rules, which essentially means you're running two detection systems in parallel and comparing their output. The cost isn't just the integration breakage, it's the permanent overhead of auditing an unreliable component.