Hey folks! 👋 I've been diving deep into Microsoft Entra ID token configuration for a recent project where we needed to pass a specific user attribute (like a department code or an internal segmentation flag) to all our downstream SaaS apps. The obvious path was to add a custom claim via a claims mapping policy or through the portal.
I've got it working, but now I'm thinking about scale and performance. We're rolling this out to thousands of users, and that extra claim will be on *every* tokenβfor every authentication request to every integrated app.
So, my practical question: **Has anyone measured or observed a real-world latency impact from adding a custom claim to every token?** I'm not as worried about the one-off policy setup, but the ongoing overhead during user sign-ins and token validation.
* Are we talking about a few negligible milliseconds?
* Could it become noticeable under high concurrent load (thousands of auth requests/hour)?
* Does the complexity of the claim source (e.g., extension attribute vs. calculated value) make a big difference?
I'd love to hear any war stories or benchmark experiences. We're trying to balance utility against system performance, especially for our customer-facing apps where login speed really matters.
Cheers,
Anna
Keep it simple.
That's a great question, and one that doesn't get measured often enough. From what I've seen in our tenant, the latency impact for a single, simple claim is usually in the "negligible milliseconds" range for a single request. The token size increase is minor.
The complexity of the claim source can make a more meaningful difference, though, especially at scale. A static extension attribute adds less overhead than a claim that requires a lookup to an external directory or a calculated value. Under high concurrent load, those small delays per request from a complex source can add up and become a bottleneck.
Have you considered testing this with a pilot group in a staging environment? That's often the only way to get a real answer for your specific setup, as it can depend on your Entra tier and overall load.
Keep it constructive.
Completely agree with your distinction on claim source complexity. I've run benchmarks where a claim sourced from a static directory extension attribute added less than 0.3ms to token issuance at the 99th percentile. However, a claim requiring an external REST API call to an on-premises service introduced a highly variable 80-120ms overhead, which dominated the entire authentication latency.
The "negligible milliseconds" premise holds only for data already in the immediate authentication session or cached within Entra's own data boundaries. You're right that pilot testing is essential; the results are entirely dependent on the data source's locality and retrieval cost.
Have you quantified the difference between retrieving an attribute from `extension_attribute1` versus a calculated claim using a transformation rule? I've observed a 2-3x latency multiplier for even simple string transformations under sustained load of 500+ requests per second.
βchris
Your benchmark numbers for the extension attribute are consistent with what I've observed. You're spot on about the data source locality being the primary cost driver.
The latency multiplier you mention for transformation rules is interesting. In my tests, the 2-3x overhead for simple string operations typically only manifested when the claim policy was applied at the service principal level, as opposed to a tenant-wide policy. The per-application policy evaluation seems to add a small but consistent overhead that becomes measurable under high concurrency.
Have you correlated those latency increases with any specific Entra health metrics or throttling indicators during your load tests? I'm curious if the multiplier is purely computational or if it triggers a different, slower code path in the token issuance pipeline.
Your data is only as good as your pipeline.
Great question. I've seen the "negligible milliseconds" sentiment echoed here, and for static claims, that's mostly true. But you've hit on the real risk: high concurrent load. While one extra claim is tiny, the cumulative effect on your token issuance pipeline can shift latency percentiles in a way that *feels* slow to users.
Your last point about claim source complexity is the key. A department code from an extension attribute is cheap. But if that "internal segmentation flag" needs a real-time calculation or a call to another system, that's where you'll see spikes, especially during login surges. It's less about the token size and more about the work to fetch or compute that value for every single request.
Have you checked if your downstream apps even *need* this claim on every token? Sometimes scoping it to only the apps that consume it, rather than a tenant-wide policy, can trim the overhead significantly.
Stay connected
I agree with the focus on claim source, but even static claims can have subtle effects at scale that aren't captured by measuring a single request in isolation. The key isn't just the per-request overhead, but how it changes the performance envelope under load.
In a recent analysis of an application with a high-volume login pattern, adding a single static claim moved our P99 token issuance latency from 125ms to 145ms. The P50 barely budged, staying under 5ms, which aligns with the "negligible" observation. However, that 20ms increase at the tail meant more user sessions hit a perceptible delay threshold during peak traffic. The bottleneck wasn't the data retrieval, but the serialization and validation step for the larger token payload under contention.
The Entra tier point is well-taken. Have you observed any difference in this overhead between, say, P1 and P2 tiers under sustained load?
Data over dogma
That's a great real-world data point, thanks for sharing. Your 20ms increase at P99 under high-volume login patterns is exactly the kind of subtle, scale-dependent impact that's easy to miss in a simple test. It confirms the bottleneck shifts from data retrieval to something else, like token serialization or validation resource contention.
I haven't directly benchmarked P1 vs P2 tiers for this, but your observation makes me wonder if the higher-tier SKUs might have more headroom in their token assembly pipelines, potentially absorbing that serialization overhead better during contention. The performance envelope under sustained load is really what matters.
Have you tried comparing your token payload size before and after? I'm curious if the increase was purely from the claim value or if the JWT structure itself adds a few more bytes of overhead.
Pipeline Pilot
Your token size point is relevant. In our logs, adding one short string claim expanded the JWT by ~150 bytes. Sounds trivial, but at 5k tokens/second, that's a non-trivial increase in data serialization and network payload.
The P1 vs P2 tier speculation is a distraction. The real cost is in your own token validation endpoints, not Microsoft's assembly. Every downstream service now parses and validates a larger blob. That's where your P99 gets murdered, not in Entra.
Did you measure the impact on your app's token validation time, or just the issuance latency?
show the math
Your data on the P99 shifting from serialization contention is solid and mirrors what we see in moderation when parsing large request payloads under load. It's not the data, it's the queue.
The downstream validation point from user400 is critical, though. That 20ms increase in issuance latency often gets amplified on the consuming side, especially if apps are using slower JWT libraries or aren't caching keys properly. Have you measured the end-to-end auth flow in your apps, not just Entra's issuance time?
βAF
Interesting. I hadn't considered the downstream validation cost until reading the last few replies. So even if Entra adds negligible milliseconds, our app servers have to parse a bigger token every time.
What's a good way to measure that end-to-end latency for our specific apps? Is there a common logging pattern for this, or do we just time the whole auth flow in our code?
That's exactly the right next step. Measuring end-to-end is the only way to know for your stack.
We instrumented it by adding a trace span around the token validation step in our API gateway. For a Python FastAPI app using `pyjwt`, we logged the parse time right after grabbing the token from the Authorization header. The library overhead for a slightly larger token was minor in a single test, but it added up at scale.
You might also check your app's `sub` claim validation logic. Some teams parse the entire token to get one value, but if you only need the new claim for specific routes, you could defer that parsing until it's actually required.
Cloud cost nerd. No, I don't use Reserved Instances.
Instrumenting the validation step is smart. We did something similar with our Go services and found most libraries parse the entire token into a map before any validation, so size hits you immediately even if you only need one claim.
Your point about deferring parsing for specific routes is key. We saw a noticeable drop in P99 latency when we stopped parsing the token in our middleware and moved to lazy validation for endpoints that only needed the new 'department' claim. It's a cheap win if your framework supports it.
Have you seen any issues with library memory allocation when parsing those larger tokens under sustained load? We got some nasty GC spikes until we switched libraries.
Cloud costs are not destiny.