Correlating timestamps from two separate systems is a forensic nightmare in a real incident. You're assuming both logging pipelines are up, clock sync is perfect, and you can query them fast enough under pressure. It's not a solution, it's a massive data reconciliation project you're now on the hook for.
And vault integration for per-host keys just replaces a secret on disk with a critical runtime dependency on the vault. If that fails during a scale event or config push, you're back to square one.
Show me the logs.
That's a really clean setup. I'm curious about the fail-open/fail-close controls you mentioned. Did you configure yours to fail-open during the pilot, or did you go straight to fail-close for compliance? It seems like a scary choice to get wrong.
Also, when you get the Duo prompt, does it ask for your username again, or does it know from the SSH connection?
CloudNewbie
Totally agree on the simplicity of that integration. It's one of the cleanest I've seen for layering MFA on top of an existing key-based setup. Your point about the audit trail staying intact is so important, it's often overlooked when people think about adding another factor.
I was pleasantly surprised that the Duo prompt doesn't ask for your username again, it just uses the one you already supplied for the SSH connection. It keeps the flow feeling like one continuous login instead of two separate chores.
My one added caveat from our rollout is to watch your PAM stack order on different Linux flavors. We had one CentOS host where another module interfered and caused a weird double-prompt. Took an hour of head-scratching to sort it out!
test everything twice
The username passthrough is a small detail that makes a huge difference in user experience. You're right, it feels seamless.
Your PAM stack warning is spot on. We ran into something similar on an older Ubuntu instance where the `common-auth` file had been modified for a custom sudo setup. The module order got weird and it tried to prompt for the local system password *after* the Duo push. Total confusion.
It makes you realize how much we depend on default PAM configurations being untouched. I started adding a simple `pam_stack` audit to our pre-requisite checks for any Duo rollout. Just a quick grep to see what else is in the chain. Saved us a few times.
✌️
Glad the pilot went well! That initial fleet setup you mentioned is the real hurdle. We templated our duo.conf via Ansible too, but found the secret distribution to be the tricky part. Rotating those integration keys across hundreds of hosts without a blip needs a solid rollback plan.
Also, a quick tip on that audit trail benefit: make sure your log shipper captures the auth lines from /var/log/secure or journalctl *before* the Duo module. We had a case where a network hiccup delayed the Duo event, making the logs look like a key was used without MFA for a few seconds. Took some parsing to reconcile.
customer first
That "no stored secrets" claim is a bit of marketing sleight of hand, isn't it? You've still got an integration key and secret key living on that host's filesystem. If that box gets popped, those are secrets an attacker can use to potentially approve their own MFA pushes, depending on how you've configured policies.
The fail-open control is the real gamble. Anyone setting that "for availability" during a pilot is basically admitting the MFA layer isn't mission-critical yet, which begs the question of why you're rolling it out in the first place.
And about that audit trail: how are you verifying the Duo event and the SSH key event are for the same session? You're now trusting timestamps across two systems. Good luck with that during a real incident response when every second counts.
That's a good point about the keys on the filesystem. I hadn't thought of that risk for a compromised host. Is there a common approach to mitigate it, like very short-lived keys, or is that just an accepted trade-off with the plugin model?
The fail-open question is worrying. It seems like a decision you'd have to make once and then never test. How do you ever gain confidence to switch it to fail-close?
Glad to hear it went smoothly! That one-line config change is definitely the appeal for us too.
How did you handle the initial push for users? We're worried about resistance to the extra step, even though it's just a push. Did you run into any "I already have my key" complaints?
It's funny you mention the "I already have my key" resistance, we got a ton of that initially. The key was to frame it as a control for the *server*, not an inconvenience for them. We explained that their key could still be stolen from their laptop, and this protects the shared infrastructure.
We made a short internal demo video showing the exact push flow, which calmed a lot of nerves. The biggest surprise was that the loudest complainers became our biggest advocates after a few weeks. One engineer admitted it actually saved him time, because he stopped second-guessing every login alert on his phone.
For rollout, we mandated a "buddy system" for the first login. You had to get your initial push with a teammate watching, which turned it into a quick collaborative moment instead of a solitary frustration. It also caught a few people with outdated Duo Mobile apps.
buyer beware, but buy smart
That one-line config is indeed the killer feature. But I want to emphasize the "pilot" part.
The real test isn't getting it working on a few bastion hosts. It's when you have to scale the enrollment, policy management, and support. That's where the ongoing operational cost hides. Duo's admin panel gets complex fast once you start adding granular policies for different server groups.
And about it not replacing key auth - yes, but it introduces a new single point of failure: the Duo service itself. If their push notification service has an outage, you're either locked out (fail-close) or your MFA is bypassed (fail-open). Neither is a great outcome during an incident when you desperately need shell access.
Have you run any scheduled failover tests to see what that outage scenario actually looks like for your team?
Your approach to reframing the MFA requirement as a server control is an excellent psychological tactic. We found the same resistance, and addressing the threat model of a stolen developer laptop was the only argument that consistently worked. The "buddy system" for first login is a clever, low-tech solution we hadn't considered; it likely prevented a significant volume of support tickets from one-off issues like cached credentials or app permissions.
However, that initial success with advocates can mask a scaling problem. We documented a clear pattern: engineers who became advocates were almost exclusively those working on critical, high-visibility systems where they appreciated the added audit trail. Engineers with SSH access to less-critical, internal-only dev servers often regressed to seeing it as pure overhead, especially if they log in frequently. This creates a cultural split. You eventually need a uniform policy, which means you're either over-securing some systems or creating friction for teams that don't perceive the same risk.
How did you handle that disparity in perceived value across different server tiers? Did you maintain the "buddy" approach for all environments, or was it just for the initial production rollout?
Fail-open is a business continuity decision, not just a technical one. It's about deciding what constitutes a complete loss of access versus a degraded security state.
You mention locking yourself out at 3am, but a fail-open setting during a Duo outage means you're relying solely on SSH keys. That's the exact pre-MFA state you were trying to move away from. Your audit trail will show the degraded state, but it won't stop a session. The real question is whether your team's incident response plan actually accounts for that mode and triggers a compensating control, like heightened log review for that period.
The logging you've set up is smart, but have you tested alerting on a fail-open event? If not, you might not know you're operating in that degraded state until it's too late.
—AF
You're absolutely right about SCIM being the simpler starting point for a small team. That "set it up once" promise is exactly what you need when you're wearing all the hats and can't afford a complex sync script turning into a maintenance sink.
But a heads-up on that path: make sure you understand exactly what attributes your SCIM sync is sending over. We once saw a sync that provisioned users but didn't send their group memberships, which meant none of our Duo policies applied. It looked like it was working until we realized everyone was in a default "allow all" state. The integration was green, but the security model was broken. 😬
Double-check those mappings before you call it done.
Keep it real, keep it kind.
That one-line config is the real hero, isn't it? So elegant.
I'm curious about the initial setup for a large fleet you mentioned at the end - was that the enrollment of users in the Duo admin portal, or the actual deployment of the PAM module and keys to all your bastion hosts? For the latter, we ended up using a simple Ansible role to push the config and it was a breeze. But the user enrollment... that was a different story. We had to script a few API calls to pre-provision users, otherwise it would've been a manual nightmare.
null
Glad it worked for you. That "no stored secrets" point is the real winner, but don't let it make you too comfortable. Those integration keys and secrets are still sensitive - they're just not a password equivalent. If someone snags those from a host, they can spoof auth logs or, depending on your Duo policy settings, potentially enroll new devices.
Your gripe about fleet setup is spot on. That's where the real cost hides. We scripted the key distribution with a config manager, but the "secret key" rotation for hundreds of hosts? That's a quarterly headache. You end up building a mini credential management system just for Duo.
- elle