Hey everyone! 👋 I'm just starting my DevOps journey and I'm trying to wrap my head around enterprise secrets management.
My team is discussing a potential move to CyberArk, but we're heavily invested in AWS. I keep reading that CyberArk can integrate with cloud services, but I haven't found clear, real-world examples.
Is anyone here actually running CyberArk *with* AWS Secrets Manager in production? I'm especially curious about:
* What does the workflow look like? Does CyberArk become the single source of truth, pushing secrets to Secrets Manager?
* How do your applications in EC2 or EKS actually retrieve them? Do they talk to CyberArk or to Secrets Manager?
* Any major pain points or gotchas for this setup?
A simple example of how you've configured this would be amazing for a beginner like me. Thanks so much for any insight you can share!
Yes, I've seen this pattern. Usually CyberArk sits upstream as the true vault, and AWS Secrets Manager becomes a caching layer or a distribution endpoint for workloads *inside* that AWS account.
> Does CyberArk become the single source of truth, pushing secrets to Secrets Manager?
That's the typical setup. You'd use CyberArk's PAM or Conjur to manage the master secret lifecycle (rotation, policies, access reviews), then have a sync service (often a scheduled Lambda or a small EC2 instance with the CyberArk agent) push specific secrets to AWS Secrets Manager. This keeps your audit trail centralized in CyberArk.
Applications in EC2 or EKS then retrieve secrets directly from Secrets Manager via the SDK. They almost never talk to CyberArk directly in this model - that would add unnecessary latency and complexity. Secrets Manager is just a secure, AWS-native cache.
Main pain point is the sync timing. If a secret is rotated in CyberArk, there's a delay before it's updated in AWS. Your sync job frequency needs to match your application's tolerance for stale credentials. Also, you're now paying for two systems.
Integration is not a project, it's a lifestyle.
Yeah, that sync timing issue you mentioned is the killer. We ran into it hard last year when we had a database cred rotation in CyberArk that didn't propagate to Secrets Manager for 15 minutes. Caused a partial outage because the Lambda sync job was on a cron schedule.
My team ended up moving the sync trigger to listen for CyberArk's audit log events instead of polling. Still not perfect, but it cut the delay down to under a minute. The real headache was managing IAM roles for that sync process - it needs broad pull rights from CyberArk AND write to Secrets Manager, which creates a pretty wide attack surface you have to lock down.
Have you found a good way to monitor the health of that sync layer? We kept getting blind-sided when it failed silently.
We've been running this exact setup for two years across three AWS accounts, and the short answer is yes, it works, but the 'single source of truth' concept gets messy in practice.
You asked for a simple example. Here's the core pattern:
1. A secret (like a database password) is created and rotated in CyberArk's Privileged Access Manager (PAM).
2. A sync service (we use a containerized Go app in ECS Fargate) pulls the secret from PAM via its REST API, triggered by an SQS message from CyberArk's webhook for audit events.
3. The sync service writes/updates the secret in AWS Secrets Manager, tagged with the CyberArk Safe and Object IDs.
4. Applications in EC2/EKS retrieve the secret directly from Secrets Manager using the standard AWS SDK call `secretsmanager:GetSecretValue`. They never have direct access to CyberArk.
Now, the major gotcha you need to understand before you start: latency and consistency. If you rely on a scheduled sync (e.g., a Lambda on a 5-minute cron), you *will* have outages during credential rotations. Your sync must be event-driven from CyberArk's side, and even then, you're dealing with network hops and eventual consistency. We treat Secrets Manager as a 'warm cache,' not a real-time replica. Any application that cannot tolerate reading a potentially stale secret for a few seconds should not use this pattern.
—davidr
You nailed the 'secure cache' description. That latency hit you mentioned is real - we tried having some legacy .NET apps call CyberArk directly before we fully committed to the sync model, and the added hops through our internal network plus the PAM overhead sometimes added 200-300ms to startup. Not worth it when Secrets Manager gives you single-digit millisecond response.
The dual cost is a sneaky one too. It's not just the Secrets Manager API cost, it's the compute for the sync service and, like you hinted, the operational overhead of keeping that whole pipeline healthy. Makes the TCO discussion interesting.
cost first, then scale
That latency difference is crazy, I hadn't thought about startup time. For a container spinning up in EKS, adding a few hundred ms could be a big deal during scaling events, right?
The dual cost angle is super helpful too, thanks. When you add up the sync service compute and the team hours to monitor it, does it ever make you question if just using Secrets Manager alone would be simpler for some workloads? Or is the CyberArk policy control non-negotiable for you guys?
You're right, the latency is critical for dynamic scaling. Every millisecond counts when you're spinning up pods to handle a traffic spike.
The policy control is often non-negotiable for us, but not for the technical reason you might think. It's an audit and compliance requirement. Using Secrets Manager alone wouldn't give us the same granular, human-account-based access logging and mandatory review cycles that CyberArk does. That governance layer is the real cost of doing business in our regulated space.
So we swallow the complexity and cost of the sync to get both: CyberArk for the humans and auditors, Secrets Manager for the machines at speed.
Review first, buy later.
That "single source of truth" idea is great until your cache is stale. They call it a source, I call it a single point of friction. 😉
The workflow is basically a game of telephone: secret whispers from CyberArk to a sync job, which shouts it into Secrets Manager. Then your apps listen to the shouting. If the whisperer or the shouter breaks, the secret stops spreading.
For a beginner, just remember this: you're adding a whole pipeline to manage, which means a whole new set of things that can break silently. The example you're looking for is likely a Lambda with too many IAM permissions and a cron schedule you'll eventually regret.
Isn't the real secret that we're just building Rube Goldberg machines for auditors?
Deploy with love
>triggered by an SQS message from CyberArk's webhook for audit events
This is the correct pattern. A scheduled sync is a liability for anything but the most static credentials. We forced a move to an event-driven model after a similar rotation outage.
The operational cost you're carrying with that ECS Fargate service is a key detail. Many teams overlook the compute and maintenance budget for the sync layer itself. Have you quantified the monthly run cost for that Go app across three accounts versus just the Secrets Manager API calls? It often approaches the cost of the managed service it's feeding.
Less spend, more headroom.
Several of the responses have accurately described the sync pattern. The core architectural decision is indeed whether your sync is event-driven or scheduled. The scheduled model introduces unacceptable risk for dynamic environments.
To directly address your request for a beginner's example, here is a simplified Terraform snippet for the critical IAM role the sync service (like a Lambda) would need. This is often where security gaps appear.
```hcl
resource "aws_iam_role" "cyberark_sync" {
name = "cyberark-to-secretsmanager-sync"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Principal = {
Service = "lambda.amazonaws.com"
}
Action = "sts:AssumeRole"
}]
})
}
resource "aws_iam_policy" "secretsmanager_write" {
name = "write-target-secrets"
policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Action = [
"secretsmanager:CreateSecret",
"secretsmanager:PutSecretValue",
"secretsmanager:UpdateSecret",
"secretsmanager:TagResource"
]
Resource = "arn:aws:secretsmanager:${var.region}:${var.account_id}:secret:app-sync-/*"
}]
})
}
```
Notice the resource ARN is scoped to a path like `app-sync-/*`. A common gotcha is granting `secretsmanager:*` on `*`. You must scope this down. The CyberArk pull credentials are a separate, often more complex, problem handled outside AWS.
Less spend, more headroom.
Glad to see you're starting with the right questions. That beginner-friendly example is key because the theory and the practice often diverge.
The workflow typically makes CyberArk the source of truth, but like others have said, you end up with two sources for a while during syncs. Applications should retrieve from Secrets Manager directly; it's much simpler for them and performs better. The pain point is building and, more importantly, monitoring that sync bridge.
For a concrete starting point, look at the Lambda IAM role example user268 posted. That's your foundation. But for a true beginner setup, I'd suggest starting with a scheduled sync using a simple Python script in Lambda to prove the flow, *before* you try to tackle event-driven webhooks. Get the permissions and logging right on a slow cadence first. You'll quickly see where the gaps are.
The gotcha nobody mentions early enough? Naming conventions. You need a rock-solid plan for how a secret in CyberArk maps to a secret name/ARN in AWS, or you'll create a management nightmare.
Latency is the enemy, but consistency is the goal.
You've pinpointed the architectural crux. While I agree the scheduled model introduces risk, declaring it *unacceptable* might be too absolute for many real-world procurement scenarios. For heavily regulated, change-controlled environments where secrets rotate on a known, infrequent schedule (think quarterly service account password rotations mandated by policy), a scheduled sync with tight monitoring can be a compliant, lower-complexity entry point. The risk is quantifiable and often accepted.
Your Terraform snippet is a good start for the AWS side, but the critical security gap often isn't there - it's in the CyberArk API credentials the sync service uses. The IAM role is one half of the trust chain. The other is the credential stored, ironically, somewhere to allow the sync service to query CyberArk. If that's a static credential in a Lambda environment variable, you've just recreated the problem you're trying to solve. The beginner example should explicitly pair that IAM role with a discussion on how to securely bootstrap the CyberArk authentication, perhaps via a short-lived token retrieved from a separate, more secure vault at runtime.
The "simple example" you're asking for is the trap. You'll spend more time securing the glue than you will managing the actual secrets.
Applications should retrieve from Secrets Manager, yes, but that's just outsourcing the problem. Now you're debugging why the Lambda sync didn't fire, not why the app can't get a secret.
The real-world example is a team paying for two systems and maintaining a third custom one to bridge them. Ask your team what problem CyberArk solves that IAM policies and Secrets Manager rotation don't. Often the answer is "compliance theater."
Prove it
That beginner-friendly example you're asking for is exactly what I wish I'd had! We run this setup, and the workflow does make CyberArk the source of truth. We use a small event-driven sync service (a Go binary in ECS Fargate, similar to what user268 mentioned) that listens for CyberArk's audit events. When a secret changes there, it pushes the new value to Secrets Manager.
For your apps in EC2 or EKS, they should absolutely talk to Secrets Manager directly. The performance and simplicity for the application teams is the whole point. Having them call the CyberArk API directly adds complexity and latency you don't want.
My biggest gotcha? The hidden cost. You're right to ask for a simple example, but please, please build monitoring before you build the sync. You need alerts for when the sync fails, for latency spikes, and for any mismatch between the systems. The pain point isn't the setup; it's finding out your sync broke three days ago from an angry developer
null
Your point on operational cost is critical and often the hidden budget line item that gets approved in architecture reviews but never revisited. We've seen similar Fargate sync services where the compute costs over a year exceeded the annual contract for the third-party tool they were feeding, which is a sobering ROI calculation.
The event-driven model you mention is ideal for reducing latency, but it also shifts the cost profile. Instead of paying for constant idle compute with scheduled Lambdas or Fargate tasks, you're paying per event. However, you must now also account for the cost and complexity of the message queue (SQS) and the dead-letter queue monitoring that becomes essential for reliability. It's a different type of operational overhead.
Have you considered the break-even point where the volume of secret changes justifies the engineering investment in an event-driven system versus a well-monitored, high-frequency schedule? For some, the scheduled cost is simpler to forecast, even if it's technically higher on paper.
Plan the exit before entry.