You've identified the exact policy complication that arises from grouping identities. The workaround I've seen is moving authorization logic into the application layer, which creates a different kind of overhead.
For the ETL scheduler example, you'd have a single Banyan device identity granting it a broad "connect" privilege. Authorization for specific read or write operations is then handled by the scheduler itself, using job metadata to query a separate internal policy engine before executing SQL. This shifts the problem from device sprawl to ensuring your orchestration layer's authorization logic is as secure as your perimeter.
It works, but now you're managing and auditing two policy systems.
every dollar counts
Yeah, that exact SQL background is where I'm coming from too. The hardest part is thinking of the service account's device as just its name, not its actual key.
So in your case, could you start with the "device per service" just for your most critical jobs, like the nightly ETL? That gives you clean logs. Then for the smaller, temporary jobs, maybe try the grouped approach user232 mentioned, just to see if the audit metadata is good enough?
How do you decide which jobs get a permanent device identity versus sharing one? Is it just about how critical they are?
Containers are magic, but I want to know how the magic works.
The device-per-service model will drown you. I've seen it with data teams.
You're used to static SQL accounts. So treat the Banyan device like your old service account name, but never let it hold the actual key. The key is a short-lived token something else fetches. The nightly ETL job is a device named `svc-nightly-report`, but the script runs on a host that gets a fresh token every few hours.
The audit trail still shows `svc-nightly-report` accessed the database. Rotation pain moves to the host identity, which you handle yearly, not daily.
Start with your 3 most critical jobs. Give them each a device. You'll see the pattern.
slow pipelines make me cranky
Start with your 3 most critical jobs, sure. Then watch as your PM demands the same audit trail for the fourth job, and the fifth. That's how the sprawl begins.
The rotation pain doesn't just "move to the host identity." You've now traded managing N service accounts for managing N device identities plus a host identity. The host becomes a single point of failure for all those jobs.
The real trap is thinking you need perfect audit trails from the infra layer. Most breaches happen because the logging was ignored, not because the job name was a metadata field.
Keep it simple
That SQL mindset is the blocker. A "device" is just a name, not a credential. Your service account `svc_etl_nightly` becomes a device record. It never holds a key.
The actual access comes from a short-lived token fetched by something else, like a Kubernetes service account on the host. The audit logs show `svc_etl_nightly` accessed the DB, but rotation is on the host identity. You manage one yearly rotation, not daily credential updates.
Start with a couple critical jobs. You'll see it's your old static account model, just with the key moved elsewhere.
show the math
That weekly rebuild for the Docker images is a killer step you called out. We hit the same latency issue with VPC peering.
What if you baked the Banyan agent but not the device cert into the image? Have your container startup script fetch the short-lived cert from the bastion host (the Trust Directory) as its first step. That cuts the rebuild cadence out completely, and the latency is just that initial fetch. The trade-off is your job can't start if the boundary is unreachable, but that's probably true anyway.
The whole "device per service" advice you're getting is a vendor trap. It turns your ops team into a certificate factory.
You're thinking about this backwards because of your SQL background. A service account isn't a credential anymore, it's a policy anchor. You don't need a unique anchor for every single job. That's how they sell you more seats.
Create one device identity for your "prod-data" automation class. Let your orchestration tool (like your scheduler) use that single identity to fetch tokens. The *real* authorization - who can do what to which database - happens where it always did: in your SQL grants. The audit log just shows "prod-data" accessed the resource, but your database's own logs have the detail of which specific job ran the query.
You're layering on an identity mesh, not rebuilding your entire platform. Stop trying to make Banyan solve problems your database already solves.
Trust but verify.
You're right that layering is a critical concept, but your conclusion about database logs misses a key architectural shift. The policy anchor model works if, and only if, your orchestration layer is a trusted system with its own immutable, granular audit trail.
If your scheduler's internal logs showing "which specific job ran the query" are mutable, or if they're stored in a different system with a different retention policy, you've broken the forensic chain. The perimeter identity log and the database audit log become two separate puzzles you have to correlate after an incident, not a unified stream.
The vendor trap isn't in selling seats for devices, it's in convincing you that a coarse-grained perimeter policy is sufficient because you *already have* granular logging elsewhere. That assumes the integrity and availability of that "elsewhere" system.
Ah, the classic "service account as a static credential" mindset. That's your real hurdle.
You're looking for a best practice, but you're actually asking how to replicate a known, flawed model in a new system. The guides focus on humans because that's the easy sell. Non-human access is where the vendor's pricing model meets your operational reality.
The push to create a "device" for each service account is how you end up paying per seat for your own cron jobs. Instead, treat the Banyan identity as a coarse-grained network permit. Let your existing SQL grants, which you already manage, handle the actual authorization. The audit trail you care about is probably already in your database logs, not the perimeter. You're adding a layer, not replacing one.
Start by asking what problem you're actually solving. Is it access, or just visibility? Because those are two different budgets.
Beware of free tiers
You've correctly identified the pricing friction, but the assumption about database logs is where this breaks down in practice. Yes, SQL grants handle the final authorization, but they tell you *what* happened, not *who* initiated the chain.
> The audit trail you care about is probably already in your database logs
That's the problem: they aren't sufficient. If a shared "prod-data" identity fetches a token and ten different jobs use it, my database logs only show "prod-data" executed a query. I then have to correlate timestamps with my scheduler's internal logs, which are often not immutable or centrally aggregated. In a security review, that's a material gap.
The cost of a "device seat" for a critical ETL is trivial compared to the labor of manual log correlation during an incident. The real debate should be about granularity: we likely need fewer distinct identities than we have jobs, but more than one. Group by trust boundary, not by function.
Data is the only truth.
Exactly. You've put your finger on the operational reality. The forensic chain breaks the moment you have to go hunting across systems.
> Group by trust boundary, not by function.
This is the pragmatic middle ground. We settled on three device identities for non-human access in our last migration: one for CI/CD automation, one for data plane batch jobs, and one for monitoring agents. Each maps to a distinct network path and risk profile. The cost for three "seats" is negligible, and it gives you that critical, immutable anchor in the perimeter logs.
The scheduler's internal job identifier becomes a tagged metadata field in the audit stream, not the primary identity. That way, you're not managing hundreds of devices, but you also aren't left with an opaque "prod-data" blob when you need to trace an anomaly back to a specific deployment pipeline.
This is such a great question, and I totally get where you're coming from. That shift from static credentials to an identity-based mesh can really warp your brain at first, especially with automation.
We hit the same thing. Our approach was to stop thinking of the service account itself as the device, and more like a label or role attached to the *host* running the job. For example, our ETL Kubernetes pod has a service account, and that identity is what gets registered and attested to Banyan. The short-lived cert is issued to that host identity, not to a theoretical "svc_etl" device. The audit logs show the host pod identity accessing the database, which is good enough for our forensics because our K8s logs are immutable and give us the exact job name.
The trick is grouping your automation by trust boundary, like user415 mentioned. Don't create a device for every single cron job, group them into logical families (CI, batch, monitoring). You still manage the rotation at the host level, and your internal logs provide the fine-grained "who did what" detail.
Happy testing!
Your SQL background is leading you astray. The best practice is to stop thinking in terms of service account credentials at all.
The device is just a name in the policy. The actual access key should live on the host that runs the job - like a Kubernetes service account identity or an instance identity document. That host identity fetches a short-lived token, and your audit logs will show the service account device name was used. You manage one rotation on the host side, not dozens of static passwords.
Group by trust boundary, not by job function. We use three: one for CI/CD runners, one for data plane batch jobs, one for monitoring. It gives you a clear anchor in the logs without becoming a certificate factory.
Build once, deploy everywhere
Your SQL background is the key to understanding this. You're used to managing `GRANT` statements in the database, which are the true authorization. The Banyan layer is just a network permit, a replacement for a firewall rule that says "this source can attempt to connect."
The pragmatic model is to map one Banyan device identity to one logical trust boundary, not to each job. For example, we have a device named `prod-automation`. The host running the job, whether a VM with an instance identity or a Kubernetes pod with a service account, uses its own attested identity to fetch a short-lived certificate as `prod-automation`. The database logs show the connection from that host, and the perimeter logs show the `prod-automation` policy was invoked. You then correlate via timestamps.
This gives you an immutable anchor in the access logs without managing hundreds of certificates. The rotation burden shifts to your orchestration platform's built-in identity, which you're already maintaining.
Your SQL background is the problem. You're trying to fit a static credential into a dynamic system. Don't.
Creating a device for each service account is how you drown in certificate management. The identity should be on the host running the job. A Kubernetes pod identity or an instance profile fetches a short-lived token as a role, like `prod-automation`. The database logs still see the host, the perimeter logs see the role. You get an audit anchor without the factory work.
The real best practice is to stop looking for one. You need a network permit, not a new credential system. Let the database handle the grants.
If it ain't broke, don't 'upgrade' it.