I keep seeing the same pattern: teams rushing to adopt AI-powered coding agents, but treating them like a fancy linter. We're plugging these things directly into our IDEs and CI/CD pipelines, giving them access to our entire codebase, and I feel like the elephant in the room is being ignored.
The risk isn't just about today's prompt leaking to the vendor. It's about the *memory* or *context window* these agents accumulate. They're designed to learn from our code, our comments, and our commit messages to be more helpful. But what are they learning, and where is that "memory" stored or used?
Think about it. Over a few weeks, an agent could piece together:
* Internal API endpoints and their rough structure from code comments.
* Names of internal services or databases from import statements and config snippets.
* Even potential security hints from `TODO:` comments like "TODO: fix hardcoded creds here later."
This isn't abstract. If you're using a cloud-based agent, this synthesized "knowledge" could influence its model for other users, or be part of a data snapshot vulnerable to a breach. For self-hosted OSS models, the risk is about the context data persisting in a vector DB or log somewhere unexpected.
My current mitigation, which I'd love to get feedback on, is a two-layer approach:
1. **A dedicated 'sanitized' Git repo for agent use.** We use a CI job that strips sensitive patterns (e.g., specific config paths, internal domain names) and pushes a cleaned branch. The agent only gets access to that.
2. **Explicitly disabling 'long-term memory' features** where possible, and auditing any telemetry or logging settings. For example, with a local `ollama` setup, you can be strict about the system prompt:
```yaml
# Example docker-compose for an isolated run
version: '3.8'
services:
ollama:
image: ollama/ollama
container_name: dev_agent
environment:
- OLLAMA_KEEP_ALIVE=24h
volumes:
- ./sanitized_code:/app:ro # Mount read-only sanitized code
command: serve
network_mode: "bridge"
# No extra volumes for persistent model data if you want amnesia
```
This adds overhead, but feels necessary. Are others thinking about this? Are we just hoping the vendors have it covered, or is there a better framework for evaluating this specific risk? I'm particularly worried about this creeping into GitOps flows, where the agent might have access to *all* manifests, including those with sealed secrets references.