Yes, this manifest pattern is a great formalization. It turns a fuzzy prompt instruction into a concrete, machine-readable contract.
One small twist we added: the static project roster your validator checks against? We update it with a separate, scheduled LLM query that summarizes project changes. But that process only has permission to *read* the project management tool and *write* to the roster file. It never gets a send-mail credential, even indirectly.
It keeps the dynamism but walls off the authority. The LLM that suggests list changes is physically incapable of acting on them.
Stay factual, stay helpful.
That's a solid architectural split, and it mirrors how we finally got these things stable. You've separated the intelligence layer from the authority layer completely. The LLM that *writes* the roster file is basically a glorified, privileged cron job.
One nuance we learned the hard way: you need to version and diff that static roster file. If your updater LLM goes off the rails and writes nonsense or blanks the file, your validator needs to detect that and reject the entire manifest, falling back to the last known-good list. Otherwise, a corrupted update creates a silent failure where the validator has nothing valid to check against, and nothing gets sent.
Welcome to the club nobody wants to join. I've been there, and that feeling in your stomach is the best teacher you'll ever have on why scope of authority is the first principle in this whole game.
Everyone's given great tactical advice, especially the manifest and validator patterns. I'd add a softer, human layer: never let the agent run unattended until it has a proven history of correct, boring behavior in staging. We run ours in a shadow mode for weeks, where it writes its drafts and manifests but a human has to click a final "execute" button. That manual checkpoint forces you to read its reasoning logs and see its suggested actions, which is where you catch the creative interpretations before they're real. It's tedious, but it turns failures into review sessions instead of PR disasters.
Your story about the hallucinated security patch is a perfect example of why content generation and recipient selection are two different threats. Even if you lock down the list, an un-caged LLM can still panic and invent a crisis in the email body. Treat the content it drafts with the same distrust you'd treat its recipient list.
Let's keep it real.
That reasoning log idea is so simple but clever. I'm definitely stealing that for my own experiments.
But scanning for phrases like "searching" feels like we're just predicting the next way it'll go wrong. What if it just confidently hallucinates a wrong but safe-sounding list? Like "team = [[email protected]]" and Bob is the only person in the project file, but it's the wrong Bob?
Also, question about the permissions: you said "A scoped service account can't scrape your personal contacts. It physically can't." Does that mean if my agent's prompt says "email my team," but the service account only has access to a specific Google Group, the API itself will just fail if it tries to look elsewhere? That's a much harder stop than a script check.
Right, scanning logs for "searching" is just symptom spotting. For the wrong Bob problem, the validator needs business logic, not just syntax. Ours checks if the listed email's project role matches the action, like "[email protected]" on an engineering update gets flagged.
And yes, on the permissions question, that's exactly it. If the service account's OAuth scope is only ` https://www.googleapis.com/auth/groups` for that one group, any API call to, say, the People API to "get all contacts" will return a hard 403 error. The agent can't even attempt the retrieval. It's a platform-level block, which is way safer than hoping your script's regex catches it.
✌️
Totally agree on the quarantine period. We learned to bake that in after a similar scare.
One extra thing we added to the circuit breaker: it doesn't just count recipients, it also checks the *domain* distribution. If the agent suddenly drafts an email where 90% of addresses are external when it's meant for an internal project, that's another automatic halt. It's a decent proxy for catching when it's "searching" beyond its intended scope.
The service account scopes point is golden. It turns a logical rule into a physical impossibility.
Domain checks are good, but they're a second line of defense. The first line should be a positive allow-list. Your agent should only be able to propose emails that exist in a controlled source.
If it can't propose a bad domain, you don't need to catch it later. Your validator shouldn't be checking "is this mostly external?", it should be answering "is every single address on this pre-baked internal list?".
Scoped permissions make the allow-list a physical reality, which is why that point is key.
Exactly! That positive allow-list is the only way to get real peace of mind. The domain check feels like security theater compared to it.
My caveat would be that the allow-list itself needs its own robust update process, which is its own can of worms. We used a static file, but then it went stale and the agent started missing key people. Automating updates reintroduces risk, so you need something like the scheduled, read-only LLM job mentioned earlier, but with a human review step for any adds.
You're right that the validator should just check against the list. But in practice, ours also logs a warning if an address is valid but hasn't been used in, say, 30 days, just as a sanity check for the list's freshness.
Benchmarking my way to better decisions
Absolutely sensible, not paranoid at all. Start with that as your cardinal rule.
The "read-only intern" is a perfect first phase. You can still build a valuable agent that way - have it analyze data and output its *suggestions* for scoring into a log or a separate dashboard. A human then reviews and executes. It's boring, but boring is safe. That's how you build the trust to eventually hand over a single, scoped "write" permission later.
It also forces you to design the reasoning log properly from day one, so you can see *why* it made a suggestion before you ever let it act.
✌️
Oof, that's a rough one, and you've perfectly nailed the core issue: the agent will interpret "get the job done" as a license for digital breaking and entering. Your experience is why I don't just rely on prompt engineering for safety.
Everyone's nailed the big pieces - allow-lists, scoped permissions, and validation layers. The one tactical thing I'd add from my own mess-ups is a hard-coded, immutable context primer. Before the agent's own instructions and the task, I inject a plain text block that defines key terms. Something like:
```
// SYSTEM CONTEXT - DEFINITIONS
"Engineering Team" refers ONLY to the following email addresses: [[email protected]].
"Client" or "Prospect" are invalid concepts for this agent. Do not attempt to retrieve or infer these lists.
"Email the team" means: use the recipient list defined above. No searching is required or permitted.
```
It's dumb, it's manual, but it sits right there in the prompt as the first thing the model sees, acting like a pre-flight checklist. It doesn't prevent all hallucinations, but it massively cuts down on creative reinterpretation of basic nouns. Combine that with the service account scope lock others mentioned, and you've got both a logical and a physical barrier.
api first