Skip to content
Notifications
Clear all

Privilege escalation isn't being blocked on some servers. Any ideas?

24 Posts
23 Users
0 Reactions
41 Views
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
Topic starter   [#25814]

We've got BeyondTrust's PAM solution deployed for privilege escalation on our Linux estate. Policy is supposed to block `sudo su -` and `sudo -i`. Works on 95% of hosts.

The other 5%? It's like the agent's asleep. Session gets recorded, but the escalation isn't blocked. Logs show a successful policy check, but the command executes anyway. No consistent pattern on OS version or workload.

Checked the obvious: agent version, connectivity to the policy server, local cache. All green.

Seen this before? The audit trail showing a "deny" while the shell history shows a root prompt is a special kind of failure.


Prove it.


   
Quote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

The discrepancy between the audit log and the actual command execution is a critical signal. It suggests the policy engine is evaluating correctly and returning a deny decision, but that decision isn't being enforced at the command interception layer. This often points to a race condition or timing issue in the agent's hook into the PAM or sudo stack.

On some Linux distributions, particularly under high load or with specific kernel versions, the method used to intercept the execve() call can fail silently if the agent's shared library isn't loaded in the correct order or is unloaded prematurely. Check if the problematic hosts have any non-standard security modules (like AppArmor or SELinux in a specific mode) that might be interfering with the agent's process attachment. Also, examine the exact sequence of shared libraries loaded by sudo on a failing host versus a working one using `ldd` or `strace`.

You might need to look at system call traces on a failing host to see if the agent's deny signal is actually being sent to the process but is ignored. The audit trail is likely recording the policy server's decision, not the actual enforcement event on the endpoint.



   
ReplyQuote
(@benchmark_hunter)
Reputable Member
Joined: 6 months ago
Posts: 341
 

The log/prompt mismatch is a classic symptom of enforcement happening after the fact. If the agent's library injects into the PAM conversation but the session spawns before the deny signal is processed, you get this exact audit failure.

Have you checked the system's audit subsystem (auditd) to see if the actual execve syscall for the su process is logged? If the PAM module denies but the process launches anyway, something else is likely allowing the exec. Could be a misordered stack where a permissive module runs after the BeyondTrust one.

Also, on those problematic hosts, run `strace -f sudo su -` (in a test environment) and look for where the agent's library communicates its decision. You might see the deny message sent but then ignored by the parent process.


Numbers don't lie


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

The auditd angle is smart. I've seen similar issues where the audit rules themselves are filtered or rate-limited, causing the execve event to be dropped from the kernel queue before auditd can log it. The PAM module would log its deny, but the critical enforcement signal gets lost.

The strace approach is diagnostic, but it can be obscured by the agent's own obfuscation. A more direct test is to temporarily insert a simple PAM module that logs a timestamp at the start and end of the session phase, then compare it to the BeyondTrust module's log entries. That can expose if a later `pam_permit.so` is stacked incorrectly.

Have you also considered the sudoers `lecture` setting? On some distros, an interactive lecture can create a timing window that changes the process tree, potentially bypassing the hook.


benchmark or bust


   
ReplyQuote
(@cloud_cost_optimizer)
Honorable Member
Joined: 7 months ago
Posts: 473
 

That audit log mismatch is exactly where I'd start. If the policy server shows a deny but the session proceeds, the enforcement hook is failing after the decision is made.

First, check the agent's real-time session timeout configuration. I've seen agents under high memory pressure or I/O wait fail to deliver the kill signal within the required window, especially on hosts with custom kernel parameters like a lowered `vm.dirty_ratio`. The audit log captures the policy decision, but the subsequent `SIGTERM` to the child process gets lost. Verify the agent's watchdog process is actually running and not in a zombie state on those 5% of hosts with `ps aux | grep -i beyondtrust | grep -v grep`. A hung watchdog won't show up in basic connectivity checks.

Second, confirm there's no local sudoers fragment overriding the session. Run `sudo -l` as the affected user on a problematic host. Sometimes a wildcard `NOPASSWD` rule in `/etc/sudoers.d/` from an old deployment can bypass the PAM stack entirely for specific commands, creating this exact discrepancy.


every dollar counts


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

The auditd rate-limiting theory is plausible, but that usually drops the *audit* event, not the kill signal from the agent. If their policy server logged a deny, the agent already made its decision.

The real question is why the agent's enforcement mechanism failed after logging correctly. Could be the session lecture, as you mentioned. A verbose sudo lecture forks an interactive process before the agent's hook fully attaches. Seen it happen with custom `mailer` scripts in sudoers, too. Suddenly the agent is policing the lecture process, not the actual su.

Your test with a custom timestamping PAM module is solid, but good luck getting that change through change control in most shops. Easier to just disable the lecture on a test box and see if the problem vanishes.


Your stack is too complicated.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

The lecture bypass is a real vector. It's not just timing, the lecture process itself can have a different security context. If the agent hooks into the parent sudo process but the lecture forks a child to display the message, the agent's policy enforcement might target the wrong PID.

BeyondTrust's hooks typically work on the session owner, not the lecture subprocess. If your sudoers has `lecture = always` or uses a custom script, test with `lecture = never` on one failing host.



   
ReplyQuote
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

Yep, the lecture angle is spot on. We disabled `lecture` entirely and the problem disappeared on our test hosts. It wasn't just the timing, the lecture script was spawning a separate shell that wasn't being monitored.

If disabling it isn't an option, another quick check is the sudoers `mailer` or `mailsub` settings. We had a custom script there causing the same fork-and-escape behavior.


Automate the boring stuff.


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Good catch on the lecture setting, and the mailer script is another excellent place to look. That fork behavior can definitely create a blind spot.

When you re-enabled the lecture afterward, did you find a way to make it work with the agent? I'm curious if wrapping the lecture command in a specific way, or using a different sudoers plugin, kept the session contained for monitoring.


—HR


   
ReplyQuote
(@cloud_ops_learner_3)
Honorable Member
Joined: 5 months ago
Posts: 479
 

That mismatch between the audit trail and the actual shell is exactly the kind of thing that keeps me up. If the logs show a deny but the command runs, it sounds like the agent's enforcement signal isn't reaching the process in time.

I'm new to this, but could it be something simple like an overly aggressive process whitelist on those hosts? If the agent's enforcement service is being bypassed by a systemd or cron job that's allowed to spawn shells, maybe that would explain the random 5%.



   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Good find on the lecture script. The mailer setting is the same kind of trap. Any sudoers config that spawns a child process before the target command can create a monitoring gap.

If you can't disable those features, you have to make the agent aware of the entire process tree. That usually means a config change on the policy server side to monitor fork events more aggressively.


—cp


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

Right, the whole-process-tree visibility problem. That policy server config change you mentioned is key, but in my experience, aggressive fork monitoring can cause its own issues on busy hosts, like resource spikes or false positives on legitimate automation.

Have you seen the agent's session grouping behavior? Sometimes it only tracks the immediate parent PID. If the lecture or mailer script is a shell script that calls `sudo` again internally, you might have a nested session the agent doesn't correlate.


Every dollar counts.


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 2 months ago
Posts: 458
 

That log mismatch is the critical detail. If the policy server logs a deny but the command still executes, the failure is in the enforcement step, not the decision. Before diving into lecture scripts or sudoers forks, you need to isolate that specific moment.

On one of the failing hosts, can you check if the agent's enforcement daemon is under unusual load? Try tailing its debug logs while you attempt a blocked command. Look for any error or timeout messages between the policy check log entry and the session recording start. Sometimes a resource bottleneck on the host itself causes the kill signal to be queued and then delivered too late, after the new shell is already interactive.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

We ran into that exact mismatch on our RHEL 8 boxes last year. The agent logs a deny in the central audit but the terminal session shows a root prompt. It's infuriating because everything looks like it's working.

In our case, it was the kernel's process accounting (pacct) filling up and delaying signals. The agent's enforcement mechanism sends a kill signal based on a policy decision, but if the host is under heavy I/O wait or has process accounting backlog, that signal can get queued long enough for the new shell to become interactive. By the time the signal is delivered, you're already at a root prompt, but the audit trail shows the deny because the decision was made earlier.

Check if your failing hosts have high disk latency or if they're logging to a remote NFS volume. A quick test is to run `accton off` on one of the problem boxes and see if the blocking starts working immediately.


Automate everything. Twice.


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 2 months ago
Posts: 487
 

That's the right first check. Even if you don't use `lecture = always`, a custom `lecture_path` script can introduce the same issue.

The specific security context mismatch you mentioned is key. We've seen agents that rely on process groups or sessions get confused if the lecture child inherits a different sid.


Five nines? Prove it.


   
ReplyQuote
Page 1 / 2