I've been conducting a thorough evaluation of our organization's secrets management strategy, specifically focusing on the interaction between our existing CyberArk Privileged Access Manager (PAM) suite and cloud-native solutions like AWS Secrets Manager. The prevailing narrative in many DevOps circles is that services like AWS Secrets Manager or HashiCorp Vault render traditional PAM solutions redundant. However, our architecture suggests a more nuanced, hybrid approach is not only possible but operationally beneficial under specific conditions.
Our current production implementation leverages CyberArk as the system of record for human and application privileged accounts, while AWS Secrets Manager handles runtime secrets for containerized applications in EKS. The integration isn't the out-of-the-box "secrets sync" one might hope for. Instead, we've built a deliberate, one-way propagation flow using a custom operator deployed within our Kubernetes clusters. The pattern follows:
1. **Centralized Policy & Audit in CyberArk:** All secret rotation policies, compliance audit trails, and access reviews are managed within CyberArk. This satisfies our internal governance requirements.
2. **Secure Synchronization:** A lightweight, internally-developed controller (Go-based) running in a dedicated administrative EKS cluster polls CyberArk's REST API (using the Central Credential Provider) for specific, tagged application secrets.
3. **Cloud-Native Distribution:** Upon detecting a change, the controller updates the corresponding secret in AWS Secrets Manager. This makes the secret immediately available to applications via the native AWS APIs or the ASM Kubernetes Secret Store CSI driver.
```yaml
# Example of our controller's configuration for a secret mapping
secretMappings:
- cyberArkSafe: "AWS-Production"
cyberArkObject: "app-db-prod-user"
awsRegion: "us-west-2"
awsSecretName: "prod/application/rds/credentials"
rotationTrigger: "cron(@weekly)" # Policy defined in CyberArk
```
**Key observations and benchmarks from this setup:**
* **Latency:** The propagation from a rotation in CyberArk to availability in AWS Secrets Manager averages 45-90 seconds. This is acceptable for our non-immediate rotation scenarios.
* **Cost:** We incur the cost of both solutions, but CyberArk covers our on-premises and legacy systems, while AWS Secrets Manager costs are negligible for the volume of cloud-only secrets.
* **Operational Overhead:** The custom integration is a liability and requires maintenance. We've mitigated this by open-sourcing it internally and treating it as a product.
* **Security Model:** This maintains a clear separation: CyberArk governs the "source of truth" with its robust access controls, while AWS Secrets Manager acts as a high-speed, highly-available cache for AWS workloads.
My primary question to the community is whether others have pursued a similar pattern. Specifically:
* Are you using the commercial CyberArk AWS Secrets Manager connector, and if so, how does its performance and reliability compare to a custom integration?
* How do you handle secret rotation for cloud-native resources (e.g., RDS databases) where the rotation logic ideally lives in AWS Lambda, but the credential vault is CyberArk?
* Have you quantified the risk or compliance overhead of maintaining two systems versus forcing all secrets (including cloud) back into CyberArk?
I am particularly interested in any performance benchmarks, failure mode analyses, or cost breakdowns from running such a hybrid model in production at scale. The marketing materials from any vendor are, unsurprisingly, devoid of these gritty details.
—chris
This is a useful real-world counterpoint to the "one tool to rule them all" thinking that tends to dominate cloud-native discussions. The distinction you're drawing between CyberArk as the authoritative policy-and-audit layer and AWS Secrets Manager as the ephemeral runtime distribution layer is a pragmatic compromise that acknowledges the different trust boundaries and performance requirements of each environment.
I do have a question about the operational overhead of that custom operator. We've seen a few teams attempt a similar one-way propagation pattern in our own ecosystem, and the maintenance burden of that glue code often becomes a hidden cost in the TCO calculation. Specifically, how do you handle drift detection when the custom operator fails or the propagation SLO is missed? Are you relying on CyberArk's native webhook capabilities to trigger the sync, or is it purely a polling-based mechanism with a fixed interval? The risk of a stale secret in the EKS cluster without a corresponding alert is something I've seen trip up teams that otherwise have a solid architecture.
Let's keep it constructive
That policy-and-distribution separation is key. We use a similar model for Spark workloads on AWS, but with a different sync mechanism. Instead of a custom operator, we have a Lambda function triggered by CyberArk's webhooks on rotation. It updates the relevant secrets in AWS Secrets Manager, and our data pipeline IAM roles are scoped to only read from there.
The caveat we ran into is latency during high-volume secret refreshes. If you're rotating dozens of database credentials at once, that webhook-to-Lambda-to-Secrets Manager chain can introduce a slight delay. We had to build a local cache in the Spark executors with a short TTL to avoid stalling job startup. It adds complexity, but keeps the runtime dependency on AWS Secrets Manager, which is much faster for the apps.