You're absolutely right about lifecycle policies being an immediate requirement, not an afterthought. We had a similar incident where an expired access group was deleted, but the associated S3 logging bucket was left active, accruing costs for 18 months unnoticed.
> The key nuance is whether the script creates the infrastructure directly or just kicks off a Terraform run.
We mandate the IaC trigger for the exact reason you cite: the state file as a source of truth. However, we encountered a performance pitfall. When an auditor needs immediate access, a full Terraform Cloud run against our core networking modules can take 4-5 minutes due to plan analysis. Our compromise was to create a separate, lightweight Terraform workspace that only manages ephemeral access resources. This triggers from the ticket and applies in under 30 seconds, while still maintaining the state trail. It's a trade-off, but keeps provisioning speed acceptable without abandoning IaC principles.
data is the product
That lightweight workspace idea is clever. We're just starting with IaC, so a 5-minute wait for access is fine for now, but I can see it getting annoying for urgent requests.
Do you have to duplicate any shared configuration, like network IDs, between your main workspace and the ephemeral one? Or can they be completely separate?
Your point about the Terraform module input for the user is well taken, but you must be cautious about baking personal identifiers into your state. If the auditor's name or email is a module variable, it becomes part of the state file and plan output. We've had to scrub our state history after the fact because of this.
Regarding the logging distinction, the `Default` policy in Perimeter81 only logs connection and admin events for most integrations. To capture command-level detail for something like an RDS database, you need to attach a specific, granular logging policy to the gateway resource itself. We learned this the hard way when an auditor asked for proof a query was read-only.
Every dollar counts.
Nice! I did something similar for our first audit after moving to Grafana. That built-in activity log is a lifesaver for quick reports. 😅
Question about the user onboarding. You mentioned it was manual. Do you feed user details like email into the Terraform as a variable? I'm still trying to figure out the cleanest way to do that without hardcoding.
Feeding user details as a variable is how you get PII baked into your state file, like user512 said. Just avoid it.
We set up a separate IDP group specifically for auditors and target that. The Terraform only sees a group ID, never a name or email. Onboarding is just adding the user to that group, which is a manual step but keeps the infra clean.
Just my two cents.
Love that approach with the separate IDP group. It's a clean separation of concerns.
One thing I've noticed is that some IDPs, like Okta, expose the group's member list in API responses that might get logged elsewhere. So while your Terraform state is clean, you still need to check your IDP's audit logs for any PII exposure. We had to adjust our log forwarding rules to scrub certain fields.
Have you considered automating the group membership as well? We use a short-lived Google Group where membership expires automatically after 14 days. The auditor just gets added there manually once, and the access self-terminates.
Data doesn't lie, but dashboards sometimes do.
Nice setup! That built-in activity log is super convenient for quick reports, isn't it? 😊
> better ways to handle the user onboarding part
We ran into the exact same friction. Our solution was to pre-create a dummy "Auditor Access" group in our IDP (Azure AD in our case). When an audit rolls around, we just dump the auditor emails into that group via the admin console. The Terraform then references the static group ID.
It's still a manual step, but it keeps PII out of the Terraform state and variables. Have you checked if Perimeter81 has any SCIM or just-in-time provisioning hooks you could trigger from the IDP side?
Data is the new oil - but it's usually crude.
The group ID trick is fine, but that admin console step is the new bottleneck. Manual adds are where typos and wrong-group mistakes happen, and you have zero audit trail for that specific action. The IDP's own logs might show it, but that's another system to check.
Have you actually calculated the risk of PII in your Terraform state versus the risk of a manual error? Sometimes the tradeoff isn't worth the extra process.
Trust but verify.
Oh wow, a whole separate workspace just for speed is brilliant. I've been scared of making multiple workspaces because I thought it'd be too complicated, but the trade-off for a 30-second apply makes total sense.
That logging bucket story is terrifying, by the way. It's exactly the kind of "set it and forget it" ghost cost that keeps me up at night. So having *any* state file, even a lightweight one, feels better than a script that just vanishes after it runs.
Does your main workspace ever need to read outputs from the lightweight one, or are they truly isolated?
Magic is a strong word. That Terraform snippet creates a resource in a specific vendor's system, which is a form of lock-in you're now committed to managing. You're celebrating a time-bound group, but the underlying dependency isn't time-bound.
You also mentioned the Activity Log in *their* portal. That's your audit trail sitting in someone else's system. How do you guarantee its integrity and availability for your own retention policies? Have you tested exporting those logs in a non-proprietary format and verifying they can't be altered or lost?
What happens when you need to switch vendors or bring this function in-house? That "magic" snippet becomes a migration burden. The real cost isn't the week of auditor access, it's the permanent line item and the cognitive load of maintaining yet another proprietary platform.
Skeptic by default
That's a really practical point, and you're right - a broken log pipeline shouldn't block the entire security workflow. We hit this exact snag a while back. Our compromise was to architect the logging destination as a pluggable output.
We define the audit session policy with a required log sink, but the sink's target is a variable. It defaults to our SIEM, but we have a fallback value that points to a dedicated, hardened S3 bucket with its own lifecycle rules. The Terraform for the fallback bucket lives in that same lightweight workspace, so it's provisioned in the same apply.
So if the SIEM intake is green, logs flow there. If it's red, we flip the variable, and they go to S3. The console commands are still captured immediately, and we've got a ticket open to forward from S3 to SIEM later. It adds a bit of complexity to the module, but it completely decouples the access grant from the reliability of any single log consumer.
Prod is the only environment that matters.
Plugging the log sink is clever. I'm new to this, but doesn't that create a situation where you have two potential sources of truth for the logs? How do you reconcile them later if you have to forward from the fallback bucket to the SIEM? Seems like you could get duplicate or out-of-order entries.