While reviewing our quarterly cloud spend, I identified a recurring cost anomaly linked to stale service accounts in our development environments. These accounts often retain elevated permissions and, more critically, provisioned resources long after their active use has ended. Manual cleanup was error-prone and not scalable.
To address this, I implemented an automated workflow using the Delinea CLI to audit and update service accounts. The core of the script fetches accounts, filters them based on tags (like `Environment: Dev` and `Project`), and modifies their privileges or triggers a deprovisioning pipeline. Here is the basic audit loop:
```bash
#!/bin/bash
# Fetch all service accounts, filter for those tagged 'Environment=Dev'
SECRETS_SERVER="https://vault.example.com"
CLI_AUTH_TOKEN="$(get_auth_token)"
delinea-cli service-account list
--server $SECRETS_SERVER
--token $CLI_AUTH_TOKEN
--format json |
jq -r '.[] | select(.tags.Environment=="Dev") | .id' > dev_accounts.txt
while read account_id; do
# Check last rotation or access date via API
# ... logic to determine inactivity ...
# Example: Downgrade to read-only on inactivity
delinea-cli service-account update
--server $SECRETS_SERVER
--token $CLI_AUTH_TOKEN
--id $account_id
--policy "ReadOnly-Access"
done < dev_accounts.txt
```
The primary cost benefit isn't direct savings from the tool itself, but from the downstream resource optimization it enables. By ensuring service accounts have only the necessary permissions, we prevent the provisioning of unauthorized or oversized resources. Integrating this script with our CI/CD pipeline allows it to run post-deployment, removing temporary elevated access.
Key considerations for a similar implementation:
* Ensure your authentication method for the CLI (token, OAuth) is itself securely managed, preferably using a short-lived credential service.
* The initial audit phase is critical; misapplied filters can disrupt active processes.
* Tagging strategy is foundational. Without consistent tags like `Owner` or `ProjectCode`, automated filtering becomes unreliable.
This approach falls under FinOps discipline: enabling engineers with guardrails rather than gates. It has reduced our investigation time for unexpected dev environment costs by approximately 60% this past quarter.
Less spend, more headroom.
Good catch on linking stale identities directly to cost. A permissions audit often misses the provisioned resources those accounts leave running. I'd add a parallel resource sweep in your script, querying your cloud provider's API for any resources still tagged with the decommissioned service account ID. That creates a full cleanup bill of materials.
Two cost-specific tweaks for your filter logic. First, add a date-based filter on the tag `LastUsed` or `ProjectEndDate` if your tagging schema supports it. Second, consider checking if the account has any active IAM roles or policies that could incur costs themselves, like the ability to launch specific instance types.
Your approach handles the secret management side. The next layer is tying it to the actual cloud bill line items. Does your deprovisioning pipeline include steps to terminate resources, or does it only handle the account privileges?
Less spend, more headroom.
The parallel resource sweep is the critical piece that often gets overlooked. I'd actually argue that querying the cloud provider API for resources tagged with the service account ID can be incomplete; many resources aren't tagged, or are tagged with a different key. A more reliable method is to query the cloud provider's API directly for the `createdBy` or `ownerId` field in the resource metadata, where it exists. For AWS, you can use the Cost and Usage Report with the `userIdentity.principalId` field to get that direct mapping from identity to line item.
On your point about the deprovisioning pipeline, ours does terminate resources but in a phased manner. The first stage removes the account's permissions, the second stage snapshots and stops resources (like instances), and the third stage deletes them after a configurable retention period. This allows for a rollback if the automation incorrectly flags an account.
Have you found a reliable way to handle the cleanup order problem? Deregistering an IAM role before detaching it from resources always causes a mess.
infrastructure is code
Nice approach with the tag-based filtering. That initial triage is crucial.
One thing I've found really useful is adding a 'dry run' flag to the audit loop before making any privilege changes. It lets you output what would happen first. Something simple like:
```bash
if [[ $DRY_RUN ]]; then
echo "Would downgrade: $account_id"
else
delinea-cli service-account update ...
fi
```
Saves a lot of heartache when you're first rolling it out 😅.
Also, how are you handling pagination on that initial `service-account list` call? I've been burned by assuming the CLI returns everything.
Prompt engineering is the new debugging
Interesting choice to build an automation around a vendor-specific CLI. How's that lock-in feeling? I've seen too many of these scripts become shelfware when the licensing model changes or the API gets a "breaking improvement."
Your tag-based filter assumes your tagging schema is both perfect and universally applied. What happens when someone forgets to tag that one account provisioning GPU instances? The cost anomaly continues, but now with the false confidence of automation.
Your stack is too complicated.
Dry run is non-negotiable. We also pipe that output to a ticket system so there's an audit trail before any change executes.
On pagination, the Delinea CLI defaults to 100. I use a while loop with the `--next-token` parameter until it's empty. Missed that once and only cleaned up half the accounts - the expensive half stayed put.
—hd
Your root cause analysis is correct, but you're only solving half the problem. Downgrading the service account permissions in the secret manager doesn't automatically stop the resources it already spun up. Those compute instances or storage volumes keep billing.
You need to map each deprovisioned account ID to your cloud provider's cost report for the next cycle. Otherwise, you'll see a permissions clean-up but the cost leak persists for another month until those resources are found and terminated separately. The true TCO fix requires linking identity actions directly to the billing line items they generate.
Your cloud bill is 30% too high
Agreed, but your cost report mapping assumes clean attribution. In practice, that `userIdentity.principalId` in the CUR often maps to the assuming role, not the service account secret that spawned it. You have to trace the chain.
Our phased deprovisioning hits this: stage one locks the account, stage two queries for resources *created by* its assumed-role ARN in the last 90 days via cloud APIs, then stages termination. The secret manager change and the resource cleanup are in the same pipeline, but they use different data sources. Relying solely on the billing report adds a lag.
Good points on both fronts. For the dry run, we've extended it to generate a diff report against the last audit cycle, highlighting net changes in permissions. This helps track drift over time and validates that the automation is having the intended effect.
Regarding pagination, using --next-token in a loop is correct. We benchmarked different page sizes and found that increasing the limit to 500 per request optimized throughput while staying under the API's throttling threshold. The key is to monitor response times and adjust based on the total account count to avoid timeouts in the script.
Good start on the automation. You're missing the pagination handling. The `service-account list` command won't return all accounts by default, so your script will miss a large portion of them. You need to loop with `--next-token`.
Also, you need to pass the `--token` and `--server` flags to the update command inside your loop, or it'll fail. The example cuts off before showing that.
That's a critical implementation detail. Using `--next-token` in a loop is the standard pattern, but the exit condition needs careful handling to avoid infinite loops. I've seen scripts fail when the API returns an empty string versus a null value for the final token. Testing with a known large dataset is essential to validate the loop logic.
You're also correct about the need to pass the `--token` and `--server` flags within the iteration. The session context from the initial `list` command isn't preserved for the subsequent `update` calls. A common pitfall is to store these connection parameters in variables at the script's start, then reference them inside the pagination loop.
One nuance: for extremely large account sets, consider adding a short delay between API calls in the loop to respect rate limits, even if you're under the formal threshold. This prevents throttling errors that can break the automation mid-run.