I was setting up a new multi-agent workflow in AutoGen yesterday and had a bit of a scare. I configured a `UserProxyAgent` with code execution enabled, and the `AssistantAgent` almost ran a `terraform apply` that would have spun up several expensive EC2 instances I didn't need. It was a good reminder of a powerful, sometimes overlooked, safety feature.
You can enforce a human approval step for any code execution, especially for commands with potential cost or security impact. It's done by setting `human_input_mode` to "ALWAYS" on your `UserProxyAgent`. This forces the agent to ask for confirmation before running *any* code block it receives.
Here's a quick snippet from my Terraform-focused setup:
```python
from autogen import UserProxyAgent, AssistantAgent
user_proxy = UserProxyAgent(
name="UserProxy",
human_input_mode="ALWAYS", # This is the key
code_execution_config={"work_dir": "terraform"},
is_termination_msg=lambda x: x.get("content", "").rstrip().endswith("TERMINATE")
)
terraform_assistant = AssistantAgent(
name="TerraformAssistant",
system_message="You are a Terraform expert. Provide code but do not run it unless the human approves.",
llm_config={"config_list": config_list},
)
```
With this, the workflow pauses and presents the proposed code in the console, waiting for a 'y' or 'n' input. For cost control, you can get more granular by implementing custom validation logic, but starting with "ALWAYS" is a great way to avoid surprises.
It's a simple setting, but it turns AutoGen from a potential runaway train into a collaborative tool. I now use it by default for any agent that might touch infrastructure or production data. It adds a minor delay but saves so much anxiety.
-- Amy
Cloud cost nerd. No, I don't use Reserved Instances.
Setting human_input_mode to ALWAYS is a basic start, but it's a blunt instrument. You're still relying on the human to catch every expensive command in real time. That's an operational risk.
The real control should be at the permissions layer. The system should prevent the agent from even proposing a terraform apply on production without specific, pre-approved context. Otherwise you're just building a manual approval queue for a system designed to automate.
Trust, but audit.
That's a really good point about shifting control to the permissions layer. It reminds me of the principle of least privilege in traditional CI/CD. Maybe a hybrid approach is needed? Something like a "gated" mode where the system checks a predefined list of dangerous commands and only asks for human input on those, while letting everything else run. Do you know if there are any plugins or patterns emerging for that in AutoGen?
still learning
Good catch. That `human_input_mode` flag is the first line of defense.
However, it's important to benchmark the interaction cost. Every "ALWAYS" loop requires a full LLM context re-submission with the user's approval response, which can slow down iterative development significantly compared to a "NEVER" flow.
The real engineering work is in moving from blanket approval to a filtered policy. You could subclass `UserProxyAgent` to parse the proposed code block, check it against a regex list of dangerous commands (like `terraform apply` or `rm -rf`), and only then prompt the human. This keeps velocity high for safe operations while maintaining the guardrail.
benchmark or bust
While `human_input_mode="ALWAYS"` is an essential safety catch, it's critical to pair it with a deterministic system that logs and classifies these near-miss events. That scare you experienced should generate a structured alert, not just a manual memory.
Your setup is a perfect candidate for integrating a metrics pipeline. You could instrument the `UserProxyAgent` to increment a Prometheus counter (e.g., `agent_blocked_execution_attempts{command_type="terraform_apply"}`) each time a code block is presented for approval. This gives you quantitative data on how often your agent proposes high-risk actions, which is foundational for tuning the system's instructions or evaluating its economic risk profile.
Without that telemetry, you're operating blind to the frequency and context of these prompts, which makes it impossible to measure the operational burden of the manual approval loop or to justify moving towards a more granular, policy-based filtering system as others have mentioned.
>a hybrid approach is needed
That's exactly what we built for our Jenkins pipelines. We used a script approval list with a deny-by-default rule set. The agent can only trigger jobs tagged as "low-risk".
For AutoGen, you'd need to intercept the code before it hits the executor. Subclass the `UserProxyAgent` and override the `execute_code_blocks` method. Parse the first line, check against a policy file, then decide to auto-run, ask, or block.
It's not a plugin, it's a pattern. The config becomes your permission layer.
YAML all the things.
Absolutely agree on the value of telemetry. Counting those blocked attempts is the first step to understanding your agent's behavior.
I'd push it a step further and suggest sending those events to a service like Datadog or New Relic, not just a local Prometheus counter. That way, you can easily correlate them with your actual cloud spend metrics. Seeing a spike in attempted `terraform apply` commands right before a cost anomaly is much more actionable.
The integration pattern is simple: just fire a webhook from within the overridden approval method. You could even tag each event with the full agent conversation ID for traceability.
Oh yeah, that's the exact setting that saved my bacon last month. I was using AutoGen to prototype some BigQuery cost optimizations, and it almost auto-approved a query that would've scanned half a petabyte 😅
The key for us was pairing that `ALWAYS` flag with a clear system message for the AssistantAgent. We told it to explicitly state the *estimated cost or impact* of any code it proposes to run, right before asking for approval. That way, the human isn't just seeing a raw `terraform apply` - they get the "why" and the "how much" as part of the prompt. Makes the approval step way more informed.
Data doesn't lie, but dashboards sometimes do.
Oh, the performance hit from "ALWAYS" is a real concern I hadn't considered. That subclass idea for a filtered policy sounds like the right balance. I wonder if anyone's published an example of that override for `execute_code_blocks` yet? I'd love to see the pattern before trying to build it myself.
I haven't seen a complete published example for AutoGen, but the pattern is straightforward. You're basically intercepting the code string before it's executed. Here's the skeleton.
```python
class PolicyEnforcingUserProxyAgent(UserProxyAgent):
def __init__(self, *args, **kwargs):
self.blocked_patterns = [r"terraform apply", r"kubectl delete", r"rm -rf"]
super().__init__(*args, **kwargs)
def execute_code_blocks(self, code_blocks):
# Check the first block against your patterns
for pattern in self.blocked_patterns:
if re.search(pattern, code_blocks[0]):
# Trigger your human-in-the-loop here
return self.get_human_input("Blocked command detected.")
# Otherwise, proceed normally
return super().execute_code_blocks(code_blocks)
```
The real work is in making that policy list and deciding what to do. Do you just ask, or do you block entirely and tell the assistant to try a different approach? That's where your operational rules come in.
Automate everything. Twice.
That skeleton gets you 80% of the way there, but you need to consider the full `code_blocks` list, not just the first one. The agent can pack multiple commands into separate blocks in a single turn. Your loop should check them all before any execution.
Also, the `get_human_input` method doesn't return a value you can pass back directly like that; it's used to populate `self._human_input`. You'd need to structure the flow to interrupt and ask, then either continue or abort. A simpler approach is to raise a custom exception that gets caught in your main loop, triggering a manual review.
You're also missing the imports for `re`. Small thing, but it breaks the example.
Automate everything. Twice.
True, it's risky to rely on manual checks for every command. That permissions layer idea seems solid.
But as a newcomer, I'm wondering: how do you practically define that "pre-approved context" for something like a terraform apply? Is it based on the workspace name, a tag in the plan, or something else? 😅
I'd worry about making the policy so strict that the agent becomes useless for any real changes.
Good catch on the full `code_blocks` list, that's a subtle but critical detail.
>raise a custom exception that gets caught in your main loop
I like this pattern. It keeps the agent logic cleaner. You could have different exception types for 'blocked' vs 'needs review' to handle them differently in your orchestrator.
That said, you'd still want to log which specific block triggered the exception for debugging. Maybe attach the offending code snippet to the exception message?
You're right to look for a concrete pattern. While there isn't a full packaged solution I'm aware of, the subclass approach is the starting point. The critical nuance is managing state; you can't just block and return. You need to suspend execution, present the context to a human, and then resume or terminate based on their decision. This often means wrapping the agent group's interaction loop, not just the proxy agent.
For a filtered policy, I'd advise defining three tiers: automatic approval for patterns matching a read-only allow list, mandatory human approval for a known high-cost list, and immediate termination for destructive patterns like rm -rf. This avoids the universal lag of ALWAYS while maintaining safety rails.
—at
Exactly the three-tier approach I use for vendor contract renewals. Automatic for small service credits, manual for price increases over 5%, and immediate escalation for any auto-renew with a fee hike.
In this context, your "read-only allow list" is key. Start with a *very* short list: maybe just `terraform plan` and `kubectl get`. You can expand it later, but only after you've seen the agent use those safely in 10+ real sessions.