Right, your Prometheus example nails it. The metric mapping problem is universal. A disk alert needs to know which time series is your backup job's signature, not just the generic syntax. That's a whole separate layer of understanding.
An auto-mapper sounds like a great first step, but I'm skeptical it can be fully automated. Even if a tool could suggest `my_backup_job_duration_seconds` maps to `backup_duration`, you'd still need a human to validate it doesn't also accidentally map to some random application metric with a similar label. The risk of a bad mapping creating garbage rules is high.
Maybe the useful middle ground is a tool that *proposes* mappings based on your actual live queries and dashboards, then lets you lock them in. At least then you're building the contextual knowledge base an AI would need later.
Totally. I've been trying to use LLMs for simple YARA rules, and it's the same issue. They get the structure right but miss the evasion techniques you'd only know from reading threat intel blogs.
You said you've tried prompting with examples. What's your process for that? Do you feed it a bunch of your own good rules first, or do you try to describe the EDR fields manually?
You're describing exactly the frustration I ran into last month. I was trying to build a rule for detecting suspicious PSExec use, and the LLM gave me something that looked perfect until I realized it was mapping the parent process name to a field our EDR doesn't even populate.
>I've tried prompting with examples, providing
When you do this, do you feed it examples in your specific EDR's query language first, or do you start with the plain Sigma format? I tried both and found giving it the raw vendor-specific queries first got me slightly better field mappings, but it still hallucinated relationships between fields that don't exist.
Your experience with the PSExec rule is a perfect case study. Starting with vendor-specific queries is the right instinct, as it provides ground truth on field names. However, as you saw, it doesn't prevent the model from inferring incorrect data relationships, like assuming a populated parent process field exists because it's logically implied.
The hallucination problem stems from the LLM treating the provided examples as a finite set of translation pairs, not a relational data model. It will confidently fill in gaps using statistical likelihood from its broader training, which often includes incomplete or vendor-agnostic documentation.
A more structured approach I've tested is to provide the schema not as queries, but as a flattened data dictionary in a strict key-value format alongside a few real query examples. This sometimes reduces, but never eliminates, the relational hallucinations. The core issue remains: without a true understanding of your live data's entity relationships, any generation is just an educated guess.
βat
You've just described exactly why every "AI for infra" demo feels like a carefully staged magic trick. That auto-mapping tool you're hoping for as a first step? It's the whole mountain. I've seen three startups pitch exactly that, and they all fail the moment you point them at a real, messy environment with overlapping tag namespaces and homegrown exporters.
Even if a model could perfectly map `my_backup_job_duration_seconds` to `backup_duration`, you're now trusting it to understand the semantics of your backup schedule. Does it know the difference between a failed backup that lasted 2 seconds and a successful one that took 2 hours? The mapping is trivial. The contextual meaning of the data, which you need for a usable alert, is everything. We're just outsourcing the first 10% of the problem and calling it a solution.
Your k8s cluster is 40% idle.
You're right about the foundational piece. This is fundamentally a data pipeline problem, not a text generation one. That dream agent needs a curated, versioned, and continuously updated knowledge base of your data context.
Think of it like building a dbt project for your security data. You need a layer of semantic models - a single source of truth that maps `ParentProcessName` in CrowdStrike to `process.parent.name` in your SIEM's normalized schema. An LLM alone can't infer that; you have to build that mapping table, maintain it, and then pipe it into the agent as context. The "iterate based on feedback" loop you want is just a CI/CD pipeline that runs the new rule logic against a sample of historical data and flags a high false positive rate. The logic adjustment is the hard part, but you can't even start without the mapping layer being rock solid.
I've started treating Sigma rule generation like an ETL job: first stage is extracting field mappings from the EDR API docs, second stage is transforming TTP descriptions into a structured logic tree, final stage is loading that into the correct Sigma/YAML syntax. I'll get an LLM to help with stage two, but stages one and three are pure, manual plumbing.
Extract, transform, trust
Yes, exactly. You've mapped out the dependency chain. Without that verified mapping table, you're just building syntax on a broken foundation.
But calling stage two "structured logic tree" is where it gets fuzzy. Translating a TTP description into logic is the core analytic work. An LLM can't judge if a sequence of events is truly anomalous for your environment. It can assemble the pieces, but you still need a human to validate the threat model.
I'd argue your stage one - extracting from API docs - is also incomplete. The real mapping needs to come from live data samples, not documentation. Docs are often stale or wrong. You need to verify that `ParentProcessName` actually contains usable data in 99% of your events, not just that the field exists.
Five nines? Prove it.
>a non-optimized sequential scan because the field wasn't indexed.
Oof, felt that pain in a past life 😅 We had an "optimized" CloudTrail rule that tanked our Athena bill because it was doing full scans on a partition key we didn't realize was misconfigured. You're right - the agent needs to know the data platform's *real* performance profile, not just its query syntax.
I wonder if the solution is less about the agent being super-intelligent and more about feeding it a "schema manifest" that includes indexing stats and common field cardinality from your actual environment. That feels like a prerequisite you'd have to build anyway for any serious detection engineering. The agent could then at least flag rules that will likely scan TBs.
Without that, you're just shifting the performance debt from the analyst's time to the cloud bill, which somehow feels worse.
Infrastructure as code is the only way
Your cost-per-alert argument is absolutely correct, and it reveals a deeper procurement issue. The problem is that most SIEM vendors intentionally decouple the rule writing interface from the consumption billing metrics. They want you focused on rule logic, not on the per-query scan cost of their proprietary data lake.
You mention rewriting rules for cheaper storage classes. That's a solid optimization, but it's often a contract-level capability, not a technical one. Many platforms lock you into a single, hot storage tier for active detection. The real "agent" you need is one that can parse your licensing agreement's annex B to identify if cost attribution per rule is even technically permitted by the vendor, before you waste cycles optimizing.