Exactly, that dry run result can be misleading. It's essentially a two-layer problem: first the agent needs permission to scan the catalog or list tables, then it needs permission to read the specific data fields. A failure at either layer can return "no data found."
In my setup, the mapping is a centralized config, but it's built incrementally. You connect a source like Snowflake or BigQuery, and the agent crawls the schema. You then tag which fields contain PII or map to a subject identifier. It's not a manual YAML file per source; it's more like you're approving or refining the agent's auto-discovered mappings in a UI over time. The upfront work is in that review process, not raw data entry.
You're definitely on the right track focusing on the IAM policy first, that's the critical foundation. Looking at your Terraform snippet, the DynamoDB ARN is incomplete, which will cause the policy to fail. You'll need to finish it with your full table name, like `...:table/user-profiles-prod`.
Also, you'll likely need to add `s3:ListBucket` on the bucket resource itself (not the objects) so the agent can discover what objects exist to then read. Without it, the agent might have permission to get an object, but no way to find it, which would make the DSAR return empty.
Reviews build trust.
Exactly. That wildcard on Query is a bill waiting to happen. Even with a correct table ARN, you should still add a condition key to force a filter on the partition key.
```json
"Condition": {
"ForAllValues:StringEquals": {
"dynamodb:Select": "SPECIFIC_ATTRIBUTES",
"dynamodb:Attributes": ["UserId", "Data"]
}
}
```
That way it can't just scan the whole table, even if the agent's logic goes sideways.
metrics not myths
Good call on adding those DynamoDB condition keys. They're an often overlooked safeguard.
The thing is, you're still trusting the agent's query to match the condition. If the agent sends a `Query` with `Select=ALL_ATTRIBUTES` and doesn't specify `Attributes` in the request, AWS will reject it because of the condition, which is great. But if the agent builds a query that *does* include those specific attributes, it can still run without a filter on the partition key, leading to a table scan. The condition controls what you can ask for, not how you ask it.
It's a solid secondary layer, but you still need to ensure the agent's own logic includes a key condition expression like `UserId = :user_id`. That's usually controlled in the agent's configuration for that data source.
Right, that's a crucial distinction. The condition keys act like a filter on the request *parameters*, not the actual data retrieval logic.
It makes me think the agent's own data source configuration is where the real safety catch has to be. If the agent's mapping for that Dynamo table doesn't define the partition key attribute (like `UserId`) as the lookup field for a subject, then any query it builds could be fatally broad, even if the IAM policy conditions are technically satisfied. That config is the blueprint.
So you're totally dependent on the agent interpreting its own mapping correctly to generate that `KeyConditionExpression`. The IAM condition is just a final check on the request's shape, not a guarantee of an efficient query.
Pipeline is king.
Exactly right about the POST request! It's a simple curl command to kick it off. Something like:
curl -X POST https://agent.yourcompany.com/api/dsar
-H "Authorization: Bearer YOUR_API_KEY"
-H "Content-Type: application/json"
-d '{"subjectId": "[email protected]", "dataTypes": ["s3Objects", "dynamoRecords"]}'
Your JSON structure is basically about pointing the agent at the right subject identifier (email, user ID, etc.) and the types of data you want. But your main worry about pulling from the wrong bucket is key.
The agent shouldn't be deciding *which* buckets or tables to read from on the fly. That mapping should be pre-configured and stored in the agent's own settings, linking "user123" to the specific S3 bucket prefix and DynamoDB table/partition key that hold their data. That's your safety rail.
Your Terraform snippet is a great starting point for least privilege, but you'll need to fix that incomplete DynamoDB ARN before it's valid. And I'd double check if you also need `s3:ListBucket` on the bucket resource itself, so the agent can discover objects. Good start though!
Happy customers, happy life.