Skip to content
Notifications
Clear all

What is the best way to handle service accounts and non-human access?

72 Posts
67 Users
0 Reactions
250 Views
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Seen that exact path. The hidden cost is the operational SLO.

Your "simple proxy" needs:
- 99.9% uptime
- sub-second p95 latency
- automated audit trail
- on-call rotation

If you can't commit to building that as a core platform service, you're just creating a time bomb. It's not about avoiding a vendor ticket, it's about whether your team can own a critical auth pipeline.


Benchmarks or bust.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Great question, and I'm wrestling with the same shift from static credentials. In our email automation setup, we ended up creating a separate Banyan "device" for each service account, because we needed clear audit trails per job. But the rotation piece feels heavy.

How are you handling the audit requirement they mention? Is it worth the overhead of a device per account versus a shared identity?



   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

You're right that most Banyan documentation centers on human users. Our team took the device-per-account route for initial audit clarity, but it created a scaling issue. The rotation overhead became significant as we passed a couple dozen service accounts.

We found API tokens for the Banyan CLI useful for automated provisioning scripts, but they still tie back to a human's identity in the audit logs. That muddies the water if you need to trace an action specifically to a job. How are you planning to attribute actions in your logs? Is a distinct identity for each automated process a hard requirement, or could you group some by function?



   
ReplyQuote
(@helenb)
Estimable Member
Joined: 3 months ago
Posts: 128
 

Coming from SQL service accounts, the device-per-account model seems logical for audit clarity, but the rotation becomes a real burden fast. How many distinct automated processes do you have? We grouped ours by risk tier.

We used Banyan's API tokens for lower-risk batch jobs, but you're right, the audit trail ties back to a human. For database access, we kept it as separate devices. Have you defined what a "clear audit trail" actually means for your compliance needs?



   
ReplyQuote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

Same boat coming from Salesforce automation. We treat service accounts like "robot users" with their own device in Banyan. The audit trail is clean for compliance, but rotation is manual pain.

What's your count of automated jobs? We found over 20 devices became unmanageable. Is the audit requirement from a compliance rule, or just internal tracking? That changes the approach.


Trying to figure it out.


   
ReplyQuote
(@emilyr22)
Reputable Member
Joined: 3 months ago
Posts: 229
 

You're describing a real hidden cost. I hadn't considered the full support model.

Does this mean the initial project plan should budget for an internal on-call rotation from the start, or is that level of ops commitment a sign you should just use the external vendor?



   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Exactly. Budget for on-call from day one or don't build it.

If the internal solution's failure mode is a P1 incident, you've committed to an SLO. That means funding the on-call rotations and monitoring. I've seen teams try to handwave this as "just run the proxy," then get crushed when token refresh fails during a platform outage and 200 cron jobs die.

If you can't stomach that operational tax, the vendor ticket is a feature. You're paying them to own the pager.


Sleep is for the weak


   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

Tag hygiene is the first cost center they don't mention in the sales demo. It's a manual, ongoing tax. Who's on the hook for that?

And >script a small pipeline sounds like "just build more internal tooling." That's more dev hours, more maintenance. So the real TCO is license fee plus the FTE to babysit the tags and rotation scripts.


always ask for a multi-year discount


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

Spot on. We learned this the hard way when our marketing automation platform's certs expired during a CI outage. Tying auth to a build pipeline adds a single, brittle point of failure.

Runtime fetch is definitely the way to go, even with the operational tax. It shifts the risk from a potential mass outage to a more manageable, incremental failure for services that can't get a fresh token.



   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

Runtime injection for the device profile is a sound architectural pattern, but it introduces a critical dependency on the availability of that internal API. What's your strategy for handling API unavailability? In a scenario where the API is down, do those jobs simply fail to start, or do you have a fallback mechanism like a slightly-stale-but-valid cached profile? This shifts the risk model from credential rotation failures to availability of your internal secret distribution service.

Your point about tagging is also the linchpin for operational utility. Without enforced tag conventions at provisioning time, you're correct that the audit trail collapses into noise. We mandated that the team field in the device registration request must include a JIRA ticket number linked to the service's runbook; automation rejects the request otherwise. This creates a forced linkage between identity and operational knowledge.



   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

Your example of an init container fetching the profile is a solid implementation of the pattern. It neatly avoids the static secret problem. However, I'd add a caveat on the API dependency.

>Runtime fetch... decouples it from CI/CD health

While true, it couples your service startup to your internal secret management API's availability. You need to consider its failure domain. If that API is down, does it block all new job starts globally, or do you have a locality-based deployment for it? The risk shifts from credential rotation failures to a potentially more concentrated availability problem.

On tagging, I agree it's mandatory for actionable logs. Our team enforced a naming convention and a mandatory set of tags, like `service_function` and `environment`, as part of the provisioning automation. Without that automation gate, the tag quality degrades immediately.


Plan the exit before entry.


   
ReplyQuote
(@annar)
Estimable Member
Joined: 3 months ago
Posts: 211
 

The transition from static SQL service accounts is a common challenge. The core decision hinges on your audit requirements versus operational overhead.

We adopted a hybrid model structured by the sensitivity of the target resource. For high-risk access like production databases, each service account gets its own device registration in Banyan. This provides an immutable one-to-one audit trail. For lower-risk API access from batch jobs, we use a single shared API token with heavily scoped permissions, accepting that the audit log will show that shared identity.

Your rotation question is critical. For the device-based accounts, we automated rotation via a dedicated internal service that handles the Banyan API calls. This shifts the burden from manual updates to ensuring that service's own reliability, which is a more contained problem. You must define the acceptable staleness window for a credential before a job fails, as this dictates your rotation service's SLA.


RTFM — then ask for the audit


   
ReplyQuote
(@cipher_blue)
Honorable Member
Joined: 6 months ago
Posts: 506
 

The hybrid model sounds great until you get a compliance audit. "Accepting that the audit log will show that shared identity" for lower-risk tasks is a trap.

How do you prove, definitively, which batch job in that shared pool made a specific API call at 3am? You can't. If your definition of "lower-risk" ever changes, or if an incident occurs, that audit trail is useless noise.

Also, automating rotation with another internal service just creates a new single point of failure. You've traded manual pain for a concentrated availability risk. What's the SLA on that rotation service, and how many teams are now dependent on it?



   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Great point about tag hygiene being a manual step that's easy to overlook. It's the classic observability problem - clean logs are useless without proper context. We had to build a tiny validation webhook for our service catalog that rejected device registrations missing key tags like `owner_team` or `slack_channel`. It's a small gate that saves a ton of investigative pain later.

Also, that rotation pipeline you mentioned is key. We trigger ours off our internal service catalog's state changes, not CI/CD. When a service is marked for decommissioning, it automatically revokes the device profile. It ties the credential lifecycle directly to the service lifecycle, which feels much cleaner.


Dashboards or it didn't happen.


   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

You're spot on about the centralised availability problem. We tried a similar internal API for our email platform's send credentials, and the grace period was the lifesaver.

Our caching layer stores tokens with a TTL a few hours shorter than their actual expiry. That way, if the auth API is having a bad day, most of our sending services just chug along with the cached version. We only start failing new service starts if the API is down for longer than that buffer.

But you've got to monitor that cache hit rate like a hawk. A dip is your first warning sign that the central dependency is struggling.


test everything twice


   
ReplyQuote
Page 2 / 5