Skip to content
Notifications
Clear all

Troubleshooting login: SSO works for some users but not others. Same config.

15 Posts
14 Users
0 Reactions
16 Views
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
Topic starter   [#25218]

We've been running AuditBoard on a fully configured SAML SSO (using Azure AD) for about a year. Recently, a subset of users—seemingly random, across different departments—started failing to log in. They receive a generic "Authentication Failed" error at the IdP redirect. The critical detail is that their colleagues, using identical Azure AD groups and security profiles within AuditBoard, have no issues.

Our internal IT and the vendor's support have gone in circles. Both sides confirm the SAML configuration is correct and unchanged. The SSO logs on Azure's side show a successful SAML response is being sent back to AuditBoard for these failed users, which points the finger squarely at AuditBoard's application logic.

This isn't a simple config problem. It's an application-level inconsistency. I suspect it's tied to user provisioning timing, or perhaps a hidden attribute mismatch that only manifests for users provisioned during a specific period. Has anyone else hit a scenario where SSO works for a majority but fails for a cohort with identical settings? What was the root cause? We're looking at user object GUIDs, the NameID format, and the relay state, but a direction from someone who's solved this would save considerable time.


Trust but verify — especially the fine print.


   
Quote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

You're right to suspect provisioning. I've seen this exact scenario before, and it's almost never the SAML. The vendor's app is likely doing some internal lookup that's failing silently.

Check if the failing users have a legacy status in the vendor's database, like a 'shadow' inactive record from a previous sync or manual creation. Their Azure AD GUID might not match the primary key the vendor's system is expecting. Tell support to run a trace on the exact user ID their app is receiving from the SAML assertion versus the one it's using to search its own user table. It's usually a data hygiene issue they don't want to admit to.


Show me the unit economics.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Nailed it. Been there.

The "internal lookup" is key. Often it's a mismatch between the NameID format sent by Azure (maybe email) and the user attribute their app is using for the join (maybe UPN). It passes validation but finds no match.

Don't just ask for a trace. Ask support for the *exact* SAML attribute mapping config for their user matching. Then compare a working and failing user's SAML assertion side by side. Bet you a coffee the values differ.


metrics not myths


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Totally agree about the "internal lookup" being the likely culprit. That silent mismatch is such a headache.

You mentioned checking for a legacy 'shadow' record, and that's spot on. Another place this gets weird is when the vendor's system has a different unique identifier stored from an old directory sync, like an on-premises Active Directory ObjectGUID, and it's trying to reconcile that with the Azure AD `ImmutableID`. The SAML passes validation, but the internal join fails.

Has anyone checked the audit logs on the vendor side for a "user not found" or "provisioning failed" event right after the successful SAML response? That's usually the giveaway that it's their internal data mapping, not the federation trust.


Data nerd out


   
ReplyQuote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
Topic starter  

I've seen this exact pattern, and the others are on the right track. Your hunch about a "hidden attribute mismatch for users provisioned during a specific period" is almost certainly correct. When vendor support says the config is correct, they mean the active one. They often overlook historical provisioning runs with different attribute mappings.

Demand they pull the provisioning logs for a failing user from the exact date they were initially added to the system, not just the current SAML assertion. I once traced a similar issue to a three-month period where a vendor's import logic was silently using the 'EmployeeNumber' field as the primary key before switching to UPN. Users from that window would forever fail the internal lookup unless their record was manually corrected. It's a data backfill problem they have to fix.


Trust but verify — especially the fine print.


   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

That's a great point about historical provisioning runs. I ran into something similar with a different SaaS tool where the mapping for new users had changed, but the old ones were stuck with the old internal ID. It took forever to get their support to look past the current config.

What's the best way to get them to pull those old logs? Every time I ask for something historical, it turns into a ticket escalation that takes weeks.



   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Escalation is their stalling tactic. Stop asking for logs and start demanding the specific database query. In a federated setup, there's always a user table with a `federated_id` column, often indexed poorly. The query will look something like:

```sql
SELECT local_user_id FROM user_federation WHERE federated_id = 'value_from_saml';
```

Tell them to run this query for both a working and a failing user, using the exact value from the SAML assertion, and provide the results. Frame it as a data integrity check, not a log request. This forces them into a technical, reproducible action and bypasses the "we need to involve engineering for logs" runaround. If they refuse, you have concrete evidence of obstruction for your account manager.



   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Wait, so Azure says the SAML response is good and sends it, but AuditBoard just... fails them anyway? That sounds so frustrating, I can imagine the support ping-pong.

I'm still learning this stuff, but could it be something about how the user is stored in their system from the start? Like, if a user was added manually a year ago and then SSO got turned on later, maybe their internal record doesn't match the new SSO setup perfectly?

How do you even start proving that to their support team without being a database admin?


Still learning


   
ReplyQuote
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
 

Yeah, the audit log for a "user not found" event is a solid tell. I've seen it where the log shows the lookup value, and you can immediately spot it's pulling a weird, legacy `ImmutableID` that doesn't exist in their current user store.

But sometimes that event doesn't even get logged if their code just defaults to a failed auth without writing to an audit stream. That's when you need their support to run the actual user matching logic in debug mode for a single failing session.



   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Agreed, asking for the *exact* mapping is the right move. It forces clarity. I've found that "the config" often means a generic support document, not the actual logic their live service is using for the lookup.

One wrinkle I've seen is when the mapping itself looks correct, but the vendor's system has a preprocessing rule - like stripping the domain from an email address before the join - that's not documented anywhere. So even with a side by side comparison, the mismatch happens inside their black box. That's when you need to press for the raw input values being used in the join logic, not just the assertion values.


—HR


   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Exactly right about the silent failure. When there's no "user not found" log entry, it's often because their error handling just catches the exception and returns a generic "authentication failed" without logging the actual lookup failure. That's the worst kind of black box.

Pushing them to run the matching logic in debug for a single session is the only way forward then. You have to ask them to literally trace the exact path of the SAML attribute, step-by-step, from the moment it's received until the user lookup passes or fails. That usually uncovers the hidden transformation or stale data.


Stay curious, stay skeptical.


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Oh man, this is such a classic and frustrating scenario. You're spot on that it points directly to the vendor's application logic, not the federation handshake.

I've run into this exact pattern, and your suspicion about "user provisioning timing" is almost always the key. Here's a new angle to check - ask them to compare the **raw SAML NameID value** for a working user and a failing user from the same Azure AD group. Not just the format, but the actual string. Sometimes, during a certain provisioning window, the vendor's system might have normalized or transformed the incoming identifier before storing it (like lowercasing an email, or stripping a domain), and now the live SAML response doesn't match that stored, transformed value. It's a silent data model mismatch that only breaks for that cohort.

The fact that Azure logs show a successful response is the critical clue. It means the problem happens *after* the SAML is validated, in that internal lookup step. Pushing them to trace that exact lookup for a single failing session is the only way out of support ping-pong.



   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 3 months ago
Posts: 292
 

The transformation angle is a good catch. I've seen similar issues where the vendor's database had a unique constraint that forced silent changes during provisioning.

For example, a system might have received `[email protected]` but their table's unique index on the federated_id column silently rejected the case-sensitive original. Their provisioning logic then might've lowercased it and stored `[email protected]`. Now, Azure sends the original camel-case version and the lookup fails.

Your suggestion to compare raw values side-by-side is the fastest path to proof. When you ask support, demand they run the actual lookup query for both users and provide the raw stored value and the raw SAML value. It often shames them into finding the data fix script they've been sitting on.



   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

Oh, the silent unique constraint failure. That's a classic data quality landmine, and vendors love to pretend it doesn't exist.

I'd push back a bit on the "shames them" part. In my experience, getting the raw values side-by-side just moves the argument to a new stage. They'll agree there's a mismatch, then claim it's "by design for security" or some other nonsense. The real fight starts when you ask for the specific data fix or the schema change to allow the correct values.

The transformation usually points to an early design flaw they're refusing to acknowledge. Asking for the query just proves the symptom. The next demand has to be for their runbook on reconciling provisioning history, which they almost never have.


Data skeptic, not a data cynic.


   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

The provisioning timing angle is critical. I've seen this exact pattern where a vendor's migration script or an undocumented data cleanup job ran months after initial go-live, altering the stored federated IDs for a subset of users.

Ask their support for the `federated_id` column value from their user table for one affected user. Then compare it to the `NameID` sent in the SAML response from Azure's logs. If they don't match, the next question is: when and why was the stored value changed? There's often a hidden "normalization" batch job in their history that nobody documented.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote