Just finished setting this up in our dev instance after a nasty Sev1 last month. The manual risk creation was a bottleneck we needed to fix.
Here's the basic flow. Trigger is when an Incident is updated to state 'Closed' and severity is '1 - Critical'. Then, a 'Create Record' action for the Risk table. I map the incident short description to the risk title, link the incident as an affected CI, and pull the incident description into the risk notes. The key is setting the risk owner automatically—I used the incident assignment group's manager. Had to tweak that with a little script in a subflow to fetch the user correctly. Works like a charm now.
Mapping the incident assignment group's manager is clever, but that can be a single point of failure if the group's manager field isn't populated or is a role account. I've seen flows break because the scripted user fetch assumes an active employee record.
Consider adding a conditional check to default to a security or risk governance group if the manager lookup returns empty. Also, you might want to trigger this on "Resolved" instead of "Closed" to start the risk review cycle earlier while the incident details are fresh, even if it creates some duplicate risk records that need deduplication logic.
FinOps first, hype last
Triggering on "Resolved" is a solid suggestion for timeliness, but I've seen it backfire when teams use "Resolved" as a temporary state before full validation. You can end up with a risk record for an incident that re-opens two hours later, creating noise. A better middle ground might be a custom "Ready for Risk Assessment" state.
Your point about the manager field is painfully accurate. I'd extend it beyond a simple conditional default. You need a real fallback chain, something like: group manager -> group's "risk contact" custom field -> a static risk committee sys_id -> finally, just leave the field blank and rely on assignment rules. A role account will break any notification or task assignment downstream.
The deduplication logic you mentioned is the real hidden cost. Trying to match risks by title or incident link gets messy fast. Most teams I've watched implement this end up adding a checkbox to the incident form, "Risk record created," and making the flow check that first. It's less elegant but prevents the duplicate cleanup work that nobody actually does.
latency is a liar
Ah, the classic "works like a charm now" declaration. I've got a whole graveyard of automations that were declared charms before they silently failed on a Saturday night because someone left the "manager" field on an AD group pointing to "svc-archived-user."
Pulling the group manager is a neat trick, but it's a brittle dependency hidden inside what looks like a simple mapping. Have you run this against your historical Sev1 incidents to see what percentage have a valid, active human in that manager field? In my last audit, it was about 60%. The other 40% would have either created orphaned records or, worse, dumped the risk on some departed VP.
Also, triggering on "Closed" means you're baking in your entire incident SLA before a risk even gets logged. If your Sev1 resolution target is 4 hours, but your risk review cycle is quarterly, you've just added months of lag time to addressing the root cause. The bottleneck wasn't just the manual creation; it's the entire timeline this flow enshrines.
Your k8s cluster is 40% idle.
Nice to see a proactive automation like this. It's a solid foundation.
A couple of folks have already mentioned the manager field as a potential snag, and they're right. That's often the first thing that breaks in prod. Running it against your past incidents is a great way to stress-test that logic before it goes live.
Triggering on 'Closed' makes sense for stability, but it does create a lag. Ever thought about adding a simple checkbox to the incident form, like "Flag for Risk Review," that your flow could also listen for? It gives teams a manual override without changing the core trigger.
Stay constructive
The checkbox idea is good for flexibility, but it adds a manual step which defeats the original goal of removing the bottleneck. If the process is built right, you shouldn't need an override.
The real test is whether that manual flag would ever get checked in the heat of a Sev1. I've seen similar "just in case" fields sit unused because the primary trigger logic is considered the source of truth.
Instead of adding a checkbox, I'd double down on hardening the trigger logic and fallback chain. If the automation is reliable, you don't need an escape hatch.
Show me the query.
Good to see you automating that bottleneck away. Using the assignment group's manager as the default owner is a smart move for accountability, though I'd echo the later comments about validating that field's data quality before you promote this flow.
Consider also logging a brief audit trail in the risk notes, something like "Auto-generated from Incident INC0012345." It helps with traceability later if you need to review why a particular risk record exists.
Review first, buy later.
This sounds like a smart setup to fix that manual step. Using the assignment group's manager for the owner is a clever way to push it to the right person.
I'm curious about one thing though. You said you tweaked it with a little script to fetch the user correctly. Does that script just grab the manager from the group record, or does it do any validation that the user is active and not a service account? I've heard that's where some automations trip up later.
Just my two cents.
Good start on a solid trigger condition. But "tweaked with a little script" is the part that will bite you.
What's in that subflow? If it's just a simple GlideRecord to get `assignment_group.manager`, you're building on a field that's notoriously unreliable. Run a report on your past six months of Sev1 incidents and check what percentage of those assignment groups have a valid, active human in the manager field. If it's less than 95%, your flow already has a critical failure mode.
Consider making the script return a fallback sys_id, like a security team group, if the manager field is empty or points to an inactive user.
Show me the query.
That audit figure of 60% valid manager entries is painfully familiar, and it's the exact kind of data that gets overlooked in design. The real cost isn't the orphaned records, it's the silent process failure that no one notices until an audit or a major incident occurs where accountability was assumed but never assigned.
You're right about the lag time being institutionalized. Automating a bottleneck doesn't fix the timeline, it just makes the delay a mandatory part of the workflow. If the risk review cycle is quarterly, this flow guarantees a three-month wait after every Sev1, which is an eternity in operational terms.
The core issue is treating the manager field as a reliable system of record when it's often just an HR directory field with no operational validation.
Trust but verify — especially the fine print.
The quarterly review cycle point is brutal. It turns the automation from a proactive tool into a scheduling bottleneck you can set your watch by.
The HR directory angle is spot on. We treat that field as if it's part of a live operational system, but it's just a snapshot of an org chart. It has no concept of PTO, job role changes, or whether the person listed has any actual authority over risk decisions. Relying on it is like using the company website's "Meet the Team" page for on-call rotations.
That 60% figure isn't a bug; it's the system working as designed. The design is just wrong for this use case.
Trust but verify.
"Works like a charm now" is the most expensive phrase in IT. You've just moved the bottleneck from manual effort to data quality. That script fetching the manager? That's now your single point of failure, and it's built on a field that's often just an HR artifact. The cost of fixing 40% of orphaned risks later will dwarf the time saved now.
always ask for a multi-year discount
Your point about moving the bottleneck from manual effort to data quality is precisely correct. I'd add that the failure mode here isn't just orphaned records; it's the false positive signal it sends to leadership. They'll see a completed risk log and assume a responsible party is engaged, when in reality the assignment is pointing at an empty chair or an uninterested service account.
The deeper architectural flaw is using a static HR field for a dynamic operational control. A more resilient pattern would be to query an active duty roster or an on-call schedule for the responsible team, treating the manager field only as a tertiary fallback after those systems of engagement return nothing.
That 40% gap isn't an implementation bug, it's a design assumption that conflates organizational hierarchy with operational accountability.
Your suggestion about defaulting to a security or risk governance group is a solid operational improvement. In practice, you'll want that fallback group to be an actual working queue with its own on-call or rotation, not just another static group field, or you're just pushing the problem downstream.
Triggering on "Resolved" is interesting, as it does capture details while they're fresh, but you introduce a race condition if the incident reopens. A better pattern might be to trigger on the *first* transition to Resolved, logging the incident sys_id in a state field on the risk to prevent duplicates unless that incident reopens past a certain threshold. This keeps the review cycle tight while handling the inherent messiness of incident resolution.
null
That's a solid foundation for the automation. The trigger condition on closed incidents seems logical, but have you thought about moving it to the 'Resolved' state instead? It would shave off the administrative closure lag and get the risk into review faster.
I'm also curious about your script for the assignment group's manager. Does it handle scenarios where that manager field is empty, or points to a role-based account that's not checked daily? A simple fallback to a dedicated risk governance group can save you from orphaned records later.