It's a powerful feature for housekeeping, and you're right that the platform doesn't advertise it well. The core challenge isn't technical, though, it's defining "stale" in a way that stays accurate.
I've implemented similar rules, and the key is coupling them with a metadata table to track exceptions. For example, you need to exclude specific risk states and, more critically, have a flag (like `manual_override_until`) that allows an owner to suspend the rule for a defined period. Without that, you'll inevitably downgrade something that's legitimately dormant but critical.
Data is the only truth.
That manual override flag is a lifesaver. I've seen teams add it after the first false positive, but it's far better to build it in from the start.
One thing I'd add: you need an alert for when that flag expires. Otherwise, you just push the problem out a few weeks and it auto-downgrades anyway. A simple scheduled report on `manual_override_until` dates coming due next week works wonders.
Automate the boring stuff.
Your question about building overly complex rules gets to the heart of the engineering challenge. The answer isn't to avoid automation, but to formalize the exceptions as configuration, not logic. You don't embed every "business quirk" into the rule's conditional statements.
Instead, you design the rule to read from a separate, version-controlled configuration file - like a YAML list of excluded states or a small table of risk IDs with manual overrides. The rule stays simple and auditable. The business logic, which changes frequently, lives in a declarative manifest that change management processes can review.
This approach makes the rule maintainable. When someone needs a new "Regulatory Hold" state, they submit a PR to add it to the exclusion list. You're not building a fragile monolith, you're creating a system where the policy can adapt without touching the automation engine.
infra nerd, cost hawk
Splitting logic and config is the right path. My team does this with a small DynamoDB table for overrides and exclusions. The rule itself is a simple Lambda that runs, checks the table, and acts.
But the "version-controlled config file" part is where people trip up. You've just moved the policy drift from the code to the config. That YAML file still needs governance, or you'll end up with a sprawling list of one-off exceptions no one understands.
The real trick is setting a hard limit, like "no more than 10 states on the exclusion list." If you hit that, it's a signal your rule's core logic is wrong and needs reworking, not another config band-aid.