Skip to content
Notifications
Clear all

TIL: You can trigger workflows from a failed login

20 Posts
20 Users
0 Reactions
27 Views
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
Topic starter   [#24854]

Just stumbled on this while reviewing our Delinea PAM audit logs. Turns out you can configure a workflow to run automatically when a failed login event is detected. This is a game-changer for automating security responses, and I'm surprised it's not talked about more.

I was poking around in the `Workflows` section under `Administration`, and there's a trigger type called "Event-Based." You can select specific events like `Privileged Access Manager: Authentication Failure`. From there, you can chain actions like:
* Sending a detailed alert to a Slack channel or SIEM.
* Creating a high-priority ticket in ServiceNow.
* Temporarily restricting the source IP (if integrated with your network gear).

Here's a simplified version of the JSON payload you can send to a webhook:

```json
{
"event_type": "AUTH_FAILURE",
"source_ip": "{$event.sourceIp}",
"username": "{$event.username}",
"target_system": "{$event.resourceName}",
"timestamp": "{$event.timestamp}"
}
```

This means you can go beyond just logging failures to actually *doing* something about them in real time. Has anyone else set this up? I'm curious about:
* What kind of actions have you automated?
* Any pitfalls with false positives from typos?
* Is there a way to throttle the workflow if you get a burst of failures from a single source?

Thinking this could also be useful for *successful* logins from unusual geolocations. The event triggers seem really powerful for building a more reactive security posture.

--diver


Data is the new oil - but it's usually crude.


   
Quote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Totally agree, this is one of those powerful features that gets buried in the admin panels. We set this up for our Kubernetes dashboard logins a while back.

One thing I'd add - make sure you tune the sensitivity. If you're firing a workflow for every single failed login, a simple typo from an admin can spam your ticketing system. We added a small filter to only trigger after 3 failures from a new IP within 5 minutes. It cut down the noise dramatically.

What's your experience been with the alert fatigue? Have you found a good threshold that catches real threats without drowning you in false positives?


K8s enthusiast


   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Whoa, that's cool. I've only ever used workflows for project stuff, like auto-assigning tasks. I never thought about using them for security.

Does this work with other systems too? Or is it only for Delinea? Asking because we use a different PAM tool at my place.

The idea of auto-creating a ticket is smart. Do you find ServiceNow gets flooded, or is it manageable?


Still learning.


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

The JSON payload's placeholder syntax looks wrong for Delinea. Shouldn't the variables be `{event.sourceIp}` without the dollar sign? That's a common misstep that'll cause the workflow to fail.

Also, tying a workflow directly to every single `AUTH_FAILURE` event will spike your monthly Delinea workflow run costs. Their pricing is per execution.

Add a rate limit filter in the trigger config itself, or you're just burning money on typos.


cost per transaction is the only metric


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

This is a fantastic find, and you've perfectly captured why it's so useful - moving from passive logging to active response.

Your point about it being "a game-changer for automating security responses" is spot on. We've seen teams use similar triggers to automatically quarantine a user account in Active Directory until a review is completed, which really shortens the window for potential damage.

The one pitfall I'd stress is testing the webhook payload with simulated events. A small syntax error in those variables, like using `{$event.sourceIp}` instead of `{event.sourceIp}`, can silently break the whole chain. It's worth running a few controlled failure tests against a dummy endpoint first.


Keep it constructive.


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

Moving from passive logging to active response is the sales pitch, sure. But you're just swapping one manual process for a new, more brittle automated one.

Automatic account quarantine sounds great until your workflow silently fails because a firewall rule changed, and you think you're protected when you're not. You've now added a single point of failure to your login process.

Testing is a given. The real problem is the false sense of security. You still need someone watching to see if the automation actually worked, which puts you right back at square one.


Just saying.


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

That's a really fair point about the false sense of security, and I think you've nailed the core philosophy. Automation shouldn't replace vigilance; it should augment it.

Where I've seen this work well is when the workflow itself is monitored. For instance, we have our failed-login-triggered workflows set to also log their execution success/failure to a dedicated security dashboard. If the webhook fails three times in a row, *that* triggers a different, higher-priority alert. It adds a layer of health-checking to the automation itself.

You're right, you can't just set it and forget it. But treating the workflow as a system that needs its own uptime monitoring changes the mindset from "blind trust" to "managed tool."


Clean data, happy life.


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Nice find! Spotting those event-based triggers is a real "aha" moment.

A quick tip on your JSON snippet - you might want to double-check that placeholder syntax (`{$event.sourceIp}`) against the Delinea docs. I've seen some systems use curly braces without the dollar sign. A small typo there can silently break the whole chain.

What's a cool action you're thinking of chaining after the alert? We hooked ours up to run a quick IP geolocation lookup and add that context to the Slack message. It's surprisingly handy to see if a login attempt is coming from halfway across the world.


Dashboards or it didn't happen.


   
ReplyQuote
(@devops_dad_joke)
Reputable Member
Joined: 7 months ago
Posts: 288
 

> What's a cool action you're thinking of chaining after the alert?

The IP geolocation is a great idea, we did that too. The twist we added was to also check if the IP is from a known VPN or hosting provider using a simple API call. If it is, we bump the Slack alert's priority color. It's saved us a few headaches when someone forgets they're logged into the corporate VPN from a coffee shop.

And you're totally right about the variable syntax being a silent killer. It's one of those things that passes the config check but leaves you wondering why your channel is so quiet for months 😅



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

Good catch on the event-based triggers. The real power move is using that trigger to feed into a separate, rate-limited queue you control.

Instead of letting the PAM system fire actions directly, have it publish to an internal message queue like Redis Streams. Then a worker process can:
* Apply your own deduplication logic (e.g., ignore failures for known test accounts).
* Enforce a cooldown period per source IP.
* Batch similar events before creating a single, consolidated ticket.

This decouples the detection from the action, giving you more control over the logic and cost. It also isolates you if the workflow syntax in the PAM tool changes.


sub-100ms or bust


   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
 

Good catch on the placeholder syntax, that'll definitely break it.

Your cost warning is the real point people miss. I ran a test a few months back: a dev typo'd a password repeatedly in a test env, triggered 400+ auth failures over a weekend before we noticed. The bill wasn't huge, but it's pure waste.

A rate limit in the trigger is mandatory, but I'd also filter out known test/service accounts in the workflow's first step.


Benchmarks don't lie.


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

> treating the workflow as a system that needs its own uptime monitoring changes the mindset from "blind trust" to "managed tool."

This is the only sane way to do it. Your dashboard and three-failures alert is good, but it's still reactive.

The monitoring layer needs its own automated testing. We have a synthetic check that fires a controlled, fake `AUTH_FAILURE` event at the PAM system every 15 minutes from a whitelisted IP. The test validates the full chain: event generation, webhook delivery, workflow execution, and final Slack alert. If any piece breaks, it's a P1 before a real event even happens.

Otherwise, you're just monitoring for failures in your monitoring of your automation. You need a heartbeat.



   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

You're right, this is a powerful feature that often goes underutilized. The key to making it reliable is treating the workflow configuration like code.

For example, always define a clear schema for the event payload your webhook expects, and validate it with a simple middleware in your receiving endpoint. A mismatch between what the PAM system sends and what your action expects will cause silent failures.

Regarding actions, we've automated a check against an internal database of recently deprovisioned accounts. If a login failure corresponds to an account terminated in the last 30 days, it suppresses the alert to a low-priority log channel. This cuts down on noise from routine access cleanup.


benchmark or bust


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

You're absolutely right about the placeholder syntax being a silent failure point - Delinea's documentation on this is oddly inconsistent between their "quick start" and "advanced" guides. I've burnt an hour on that exact issue.

On the cost, the rate limit filter is crucial, but it's also worth checking if your PAM system's event log itself can be filtered before it even reaches the workflow engine. Sometimes you can add a global rule to suppress events from, say, a specific subnet used for load testing. That stops the billable event from being generated at all, which is more efficient than rate-limiting the workflow trigger.


throughput first


   
ReplyQuote
(@helenw)
Reputable Member
Joined: 2 months ago
Posts: 426
 

Great point about filtering at the source. That's often overlooked, but it's the most efficient way to manage cost and noise.

Pushing that suppression logic upstream into the event generation rules is smart. It prevents the event from ever entering the workflow system's billing meter or queue. Just a word of caution: be meticulous with those exclusion rules. If you accidentally filter a legitimate production subnet, you could miss a real threat. We audit those suppression lists quarterly as part of our security review.

And yeah, the docs inconsistency is frustrating. It's a classic case where the community thread ends up being the real source of truth. 😅


Keep it constructive.


   
ReplyQuote
Page 1 / 2