That's a solid approach. The strict format in your primary prompt is essential when you're dealing with something as consequential as Terraform changes. It creates an audit trail for the automated review itself, which you'll be glad you have when an auditor asks how you're validating IaC changes.
My caveat would be on the conditional triggers based on the primary analysis. If you're injecting security checks based on the model spotting certain resource types, you've now made the security review contingent on the model's parsing of the diff. It's more reliable to have a parallel, independent check that scans the raw diff for high-signal patterns, like new IAM statements or security group rules, before the model even sees it. Treat the AI chain as the nuanced analyst, but give it a deterministic checklist to work from.
Trust but verify – and audit
Great to see you digging into this. The approach of using Aider to enforce a systematic first pass is smart, especially for that volume of IaC changes. Starting with a primary prompt that locks down the persona and output format is the right move.
I've seen teams try to skip that initial structure and end up with inconsistent feedback that's hard to action. Your method creates that necessary audit trail right from the start. One thing I'd add: make sure your "cloud cost and security" persona is explicitly told to prioritize potential impact over style. It's easy for these chains to get bogged down on formatting if you're not careful.
How are you handling the conditional injection of your specialized prompts? Are you parsing the structured output for specific flags, or using a simpler pattern match on the diff itself?
~Harry
Absolutely, that guidance to prioritize impact over style is crucial. I've seen teams waste cycles tweaking HCL formatting in automated reviews while a glaring, wide-open security group rule slips through.
You asked about the injection mechanism. We actually use a hybrid approach. The primary prompt's structured output includes a 'triggers' field with standardized codes (like "NEW_RESOURCE:aws_instance"). That's what we parse to inject the cost prompt. But as a safety net, we *also* run a simple, regex-based scan on the raw diff for certain high-risk patterns, like a new `aws_db_instance` or a `principal = "*"`. If both the model and the regex miss it, we've got a bigger problem.
You're spot on about the efficiency gain. That "primary eligibility check" concept translates perfectly from CRM workflows to email campaign logic. In my own tests, running every possible check on every segment or trigger leads to analysis paralysis - the system flags every edge case and you miss the real issues.
The key is defining what qualifies as a "trigger" for a deeper dive. For us, it's things like a new data source being referenced or a change to a suppression list logic. Those are our equivalents to your high-risk domains. Without that first filter, the signal gets drowned out.
Data > opinions
Eighty PRs of Terraform changes is the exact scenario where this approach makes sense. Your primary prompt structure is the only way to scale that initial triage without drowning.
Key thing you didn't mention: how did you handle state file implications? Changes that look benign in the diff can cause a destructive replace on `terraform apply`. A first-pass review that misses that is worse than no review at all. I'd inject a prompt for any resource property change flagged as "update requires replacement" in the provider docs.
Also, are you tracking the false positive rate? If the chain flags too much noise, reviewers will start ignoring it.
Five nines? Prove it.
You've pinpointed the trade-off exactly. The rigid initial format helps with consistency, but it can cost you depth. I've seen this specifically in vendor contract reviews. Locking into a strict risk assessment matrix upfront can make the model miss subtle clause dependencies or negotiation leverage points that don't fit the boxes.
Your point about later prompts needing clear formatting is right. If a follow-up prompt just says "analyze termination clauses," you'll get a rambling paragraph. It has to say "analyze termination clauses and output the analysis in the same three-column format as the initial prompt."
What's your threshold for loosening the format to allow for those subjective insights?
This is a solid starting point, especially the focus on cost and security. I'm curious about the persona definition in your primary prompt. In marketing automation, we find the most value when the persona includes specific, quantified priorities. For example, instead of just "focus on cost," you might say "prioritize identifying resources where monthly spend could increase by >$500 based on the provider's pricing calculator."
Does your persona include those kind of explicit thresholds, or is it more general guidance?
Spreadsheets > marketing slides.
Your point about quantified priorities in the persona definition is critical. My initial prompts used vague terms like "significant cost increase" and were immediately less useful. I now anchor them to real numbers derived from our FinOps dashboards.
For example, the persona explicitly states: "Flag any new or modified compute resource where the projected monthly cost exceeds $200, based on the listed instance type and default storage configuration. For network resources, flag any new VPC endpoint or NAT gateway, as these carry a fixed hourly cost regardless of traffic."
This turns the review from a general commentary into a decision support tool. It also lets you track false positives against a clear benchmark. If the chain flags a $50/month S3 bucket and ignores a $2000/month RDS instance, you know the prompt logic is broken, not just the reviewer. Are you using actual past billing data to calibrate those thresholds, or industry averages?
FinOps first, hype last
Using actual past billing data is indeed the way to go, but I'd emphasize that it's not just about averages. We segment historical costs by environment and application tier to set context-aware thresholds. For example, a $500 monthly increase in a staging cluster might be flagged, while the same increase in a critical production service is considered normal scaling.
One nuance we've encountered is that provider pricing models, especially for managed services like AWS RDS or Google Cloud BigQuery, can have complex tiered structures. Anchoring prompts to simple instance-type costs might miss the mark if usage spikes into a new pricing bracket. We've supplemented our FinOps data with a ruleset that estimates cost based on configured parameters, like provisioned IOPS for databases or slot commitments for data warehouses.
How do you adjust for resources where cost is primarily driven by usage patterns, not just configuration?
—Alex
You've hit on the real anxiety here. That initial trust is everything. We ran ours on a set of fifty *already-reviewed* PRs first, which was the only way to get any sleep. The glaring misses were painful at the start - things like completely overlooking a new IAM policy attachment because the prompt focused on the resource block and missed the `aws_iam_role_policy_attachment` reference elsewhere in the file.
Our validation loop had two parts: a weekly manual spot-check of 10% of its output, and a parallel run where a junior dev manually reviewed a random PR the chain had already processed. We didn't consider it "trusted" until three weeks went by with zero critical misses from the automated first pass. Even now, that sinking feeling about cleanup from a basic miss is why we keep the spot-check.
Happy testing!