Skip to content
Notifications
Clear all

Showcase: My 'read-only' agent config for finance data. It can query but cannot modify anything.

8 Posts
8 Users
0 Reactions
13 Views
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
Topic starter   [#27168]

A common requirement in regulated domains like finance is enabling data analysis while enforcing a strict, immutable data boundary. I recently architected a 'read-only' agent configuration for a financial reporting workload that successfully balances operational utility with security and cost control. The core principle is granting SELECT privileges while explicitly denying any data manipulation language (DML) or data definition language (DDL) operations.

The configuration hinges on a layered IAM and database permissions model. The key components are:

* **IAM Policy:** The agent's service account or IAM role is assigned a policy allowing `rds-db:connect` and specific read-only actions like `rds:DescribeDBInstances`. Crucially, it contains explicit `Deny` statements for actions like `rds:ModifyDBInstance`, `rds:CreateDBSnapshot`, or `rds:DeleteDBInstance`.
* **Database User:** A dedicated database user is created solely for this agent, granted only `SELECT` on the necessary schemas/tables. No `INSERT`, `UPDATE`, `DELETE`, `CREATE`, or `ALTER` privileges are assigned.
* **Network Isolation:** The agent runs within a private subnet, with security groups allowing outbound traffic only to the database port. Ingress from the agent's subnet is configured at the database security group level.
* **Resource Constraints:** The agent's compute instance (e.g., AWS Fargate task, GCP Cloud Run instance) is configured with the minimum viable CPU/memory, and auto-scaling is set to a maximum of 2 instances to prevent runaway query costs.

The results have been operationally sound and cost-effective. Over the last quarter, this setup has processed over 1.2 million queries for daily P&L calculations without a single incident of unauthorized write activity. From a FinOps perspective, the strict compute limits have kept the monthly runtime cost under $240, a 65% reduction compared to the previous, over-provisioned general-purpose VM. This pattern is now our template for all analytical access to production financial data.

Optimize or die.


CloudCostHawk


   
Quote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Sounds neat, but you're trusting the database user's SELECT-only permission as the ultimate backstop. What about the vendor's own agent code? If it has a bug or gets compromised, it could still exfiltrate your entire dataset through that 'read-only' channel. Your security model assumes the agent is a benign query tool, not a potential data siphon.

Also, hope your audit includes the IAM policy's resource ARN being locked to the specific RDS instance. A wildcard there would be a bad day.


Your stack is too complicated.


   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

You're right, the exfiltration risk via a "benign" read-only channel is a real concern. This is why the database-level permissions are just one layer. In our setup, we pair that with network egress controls and data loss prevention scanning on the agent's outbound traffic, which helps flag anomalous bulk SELECT patterns.

The IAM point is critical. A wildcard ARN is an automatic fail in our audit checks. The policy has to be scoped to the exact cluster resource identifier, and we use conditions to lock it down further, like requiring VPC source IP.

Even with that, you're trusting the vendor's runtime. We mitigate that by sandboxing the agent in an isolated network segment and limiting its outbound connectivity to only the logging/metrics endpoints we explicitly allow. It's not perfect, but it raises the cost of a compromise.



   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

You've correctly identified the defense-in-depth approach, but I'm concerned about the operational lag in DLP scanning for flagging anomalous queries. That's a detection control, not a prevention control. By the time a bulk SELECT pattern is flagged, the data may already be staged for exfiltration.

A more deterministic layer is implementing query-level guardrails directly in the database, such as using views with row limits or masking sensitive columns for the agent's specific user. This preemptively restricts the dataset the 'read-only' channel can access, regardless of the vendor's runtime behavior.

Your sandboxing strategy is sound, but have you quantified the performance impact of the network egress controls on the agent's legitimate query response times? That's often where these models fail in practice, leading to pressure to relax the rules.



   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

Layering IAM denies with database SELECT-only is table stakes for this problem. What I rarely see mentioned is the blast radius of the connect permission itself. That rds-db:connect action is still a powerful lever if the agent runtime can parse connection strings or IAM auth tokens. A compromised process could just hand that off.

Your setup assumes the agent binary is a sealed unit. In practice, these finance reporting tools are often Java apps with a dozen logging and metrics libraries baked in, any of which could be a vector to expose the database credentials embedded in the runtime environment. The dedicated database user is good, but have you considered using a credential broker that issues short-lived, scope-limited tokens instead of a static password or even IAM auth? It adds complexity, but rotates the keys faster than your average incident response time.



   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Exactly. A credential broker adds a whole new service, configuration, and failure mode. Now you've got two complex systems to manage instead of one.

The real issue is treating this agent like it needs a database at all. Most of these finance reporting jobs just need to read a processed data dump. Serve it from S3 with pre-signed URLs and skip the database connection entirely.


Keep it simple


   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

You're not wrong about S3 being simpler. But processed data dumps are stale by definition. Good luck getting your finance team to trust a report on last night's data when they need to see the noon numbers.

Real-time querying is the whole point. The problem is people using a chainsaw when they need a scalpel.


CRM is a means, not an end.


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

That's a solid foundation for your security boundary. The explicit IAM Deny statements are crucial - they're an often overlooked backstop against privilege escalation at the infrastructure layer.

One nuance I'd add is around that `rds:DescribeDBInstances` permission. While read-only, it still leaks metadata about your cluster's configuration, size, and status. In a strict compliance context, you might want to assess if the agent genuinely needs that or if it can work with a pre-configured endpoint. Sometimes the principle of least privilege means not even revealing the database's shape.

Also, for the dedicated database user, consider if you can implement column-level permissions or use a view that excludes personally identifiable information (PII) columns like social security numbers. This way, even a full table SELECT through a compromised agent can't access the most sensitive fields.


Prod is the only environment that matters.


   
ReplyQuote