You're spot on about the debugging nightmare shifting to the sync layer. We added monitoring after the fact and found our "event-driven" pipeline had been silently failing for two days because of a CyberArk API version bump. 😅
That said, I think "compliance theater" is a bit harsh. For us, CyberArk solves the "human in the loop" problem that IAM policies don't - the approvals, the session recording, the granular access for admins. Secrets Manager rotation is great for automated machine secrets, but can't gate a human requesting a prod DB password at 2am.
But your core warning stands: if you don't have a strong reason to need *both* systems, you're just building a fragile, expensive bridge between two islands.
edge cases matter
Several good points already. As a beginner, you're asking the right architectural question about source of truth. In our production setup, CyberArk is indeed the source, but with a crucial nuance - it's a *managed* source, not an interactive one. Applications never retrieve from it directly.
For your EC2/EKS question, the answer is always Secrets Manager. The sync pattern exists to let application teams keep using the AWS-native SDKs they already know, without needing CyberArk client libraries or convoluted IAM trust chains back to on-prem. The "simple example" trap others mentioned is real - that Lambda function becomes your most critical security component, and you'll find yourself writing more code to secure and monitor the bridge than to actually deliver secrets.
The major pain point nobody's mentioned yet is secret versioning. CyberArk and Secrets Manager handle it differently. When your sync service pushes a rotation, you need logic to handle the version staging in Secrets Manager, or you'll cause application outages during the propagation window.
infrastructure is code
Thanks for asking this, I'm in a similar spot trying to learn this stuff 😅
Everyone here is saying the apps should talk to Secrets Manager, which makes sense. But I'm still fuzzy on the actual trigger. For a beginner setup, what's a simple way to start the sync? Do you just run a cron job somewhere that checks CyberArk every hour?
Also, user1081 mentioned monitoring the sync is crucial. What are you monitoring for, exactly? Just if the Lambda fails, or something more?
Your question about the cron job is the heart of it. That's the "simple" way, and it's a trap. You now have a stale secret window you need to define and justify.
Monitoring? Don't just watch for Lambda failures. Watch for successful executions that return no changes when you know a secret rotated. Watch the timestamp delta between CyberArk's change log and the Secrets Manager update. You're not monitoring a process, you're monitoring the failure of your own architecture to be real-time.
Prove it
Absolutely agree on that final point about the credential chain, it's the trapdoor in every bridge design. Even with IAM roles, that initial CyberArk API credential becomes a permanent, highly privileged secret itself. If it's just another static entry in a config file, you haven't solved the root problem - you've just moved it and painted a target on the sync service.
One pattern we've used for the bootstrap is having the sync service (in ECS or Lambda) assume its IAM role, then use those temporary AWS credentials to fetch a short-lived CyberArk token from a dedicated, highly restricted safe. It adds steps, but it means there's no long-lived credential stored for the bridge itself. The hard part is getting that first token-granting credential provisioned securely in your deployment pipeline.
Architect first, buy later
You've hit the nail on the head with the bootstrapping problem. That initial credential provisioning is the real make-or-break moment. Your pattern is sound, but the deployment pipeline step often gets hand-waved.
We use a similar token-granting credential flow, but we had to treat that safe's credentials like a hardware root of trust. The secret is generated once during the initial platform build, placed in that safe, and then *never* exposed again. The deploy pipeline doesn't even see the raw value, it just has IAM permissions to trigger a CloudFormation custom resource that calls a secured, internal API to perform the initial fetch and inject it as a secure parameter.
The operational risk shifts from storing the credential to absolutely locking down that one safe and its API access, which feels more manageable. Have you run into issues with that initial generation and placement being a manual, break-glass procedure? That's where we've seen teams slip up.
Architect first, buy later
It can work, but the cost is a big blind spot for a new team. Everyone will sell you on the integration features, but no one talks about the extra compute and SQS costs for the sync layer.
Our budget review last quarter showed the Fargate tasks and queues cost more than the Secrets Manager usage they supported. You need to price that sync service as a permanent line item.
Also, you asked about pain points. The biggest one is when the sync breaks and you have different secrets in each system. Your apps use Secrets Manager, but your admins update CyberArk. Who's responsible for the drift? That's a process problem, not a technical one.
Hey! 👋 Welcome to the fun world of secret sprawl. The short answer is yes, this is done, but like others have hinted, the "simple example" is the siren song.
> A simple example of how you've configured this would be amazing
Okay, here's a basic Lambda sketch that gets you 80% there and 100% of the operational debt:
```python
import boto3
import requests
def lambda_handler(event, context):
# 1. Get CyberArk token (hardcoded creds in env vars, yolo)
cyberark_auth = {"username": "...", "password": "..."}
token_resp = requests.post("https://cyberark/api/auth", json=cyberark_auth)
token = token_resp.json().get("token")
# 2. Fetch secrets from specific CyberArk safe
headers = {"Authorization": token}
secrets = requests.get("https://cyberark/api/safes/myapp/secrets", headers=headers)
# 3. Push each to Secrets Manager
sm_client = boto3.client("secretsmanager")
for secret in secrets.json():
sm_client.update_secret(SecretId=secret["name"], SecretString=secret["value"])
```
This will work in a PoC. Then you'll spend the next year adding error handling, secret versioning, monitoring, alerting, and a secure bootstrapping method for that first CyberArk credential (see user1168's point). Your apps do talk to Secrets Manager, which is the only scalable part.
The gut check question: does your compliance team *require* CyberArk for audit trails on every secret, including machine ones? If not, maybe just use Secrets Manager natively.
Prompt engineering is the new debugging
> This is often where security gaps appear.
That's the part I keep getting stuck on. Everyone shows you the IAM role, but nobody shows the actual CyberArk API policy or how the safe is configured. If the lambda's CyberArk token can read from *any* safe, you've just built a master key. Is the standard practice to create a dedicated safe just for the sync service with the minimal needed permissions?
Exactly. That sketch is the recipe for a "security theater" integration. It works until your first security audit, where they'll ask how you protect that CyberArk credential in the Lambda's environment.
You mentioned adding secure bootstrapping later, but in my experience, that's where projects stall. The team gets the sync working with the hardcoded creds, then moves on to the next fire. That temporary credential becomes permanent because retrofitting a proper token service is now a "breaking change."
One compromise we used early on was to store the CyberArk API credential *in* Secrets Manager itself, locked down with a very restrictive resource policy. The Lambda's IAM role could read only that one secret. It's still not ideal, but it moves the problem into AWS's IAM/KMS boundary, which is often easier to audit and monitor than a Lambda env variable.
The simple example you're asking for is the problem. It leads to that hardcoded credential pattern everyone is showing.
Your workflow question misses the point. If you push from CyberArk to Secrets Manager, then CyberArk is the source of truth. But your apps are now dependent on a sync process you have to build, monitor, and debug.
Real pain point? Drift. Someone changes a secret in CyberArk, the sync fails silently, and your app in EKS pulls the old secret from AWS. Now you have two truths.
If it's not a retention curve, I don't care.
You're right about the hidden cost of that "third custom system." It's rarely just a Lambda. You end up with a queue for change events, a dead-letter queue, CloudWatch alarms, and a runbook for when the sync fails. It's a whole microservice to maintain, and its uptime becomes critical path for secret availability.
That's the compliance theater part - the diagram looks secure, but the operational burden and new failure modes can actually increase risk compared to a simpler, single-system approach. The question is whether the compliance requirement is about the *architecture* or the actual *security outcome*.
—Anita
The hidden cost you mentioned is the real trap, but your warning about monitoring arrives too late. Building monitoring *before* the sync service is a luxury few teams have. The business wants the integration ticking a compliance box yesterday.
Your event-driven model based on audit events is the right architecture, but you're still trusting CyberArk's event stream to be reliable and comprehensive. We had a case where a bulk update via the CLI didn't trigger the expected audit event type, and the sync missed it. You need to supplement events with a periodic reconciliation job, which adds even more to that operational cost you flagged.
And let's be honest, when that sync breaks, the "angry developer" isn't just a pain point - they're the *first* and only alert you'll get. Your monitoring is only as good as the person who remembers to update the thresholds when the secret change rate triples overnight.
Absolutely right about the angry developer being the first alert. It's a clear signal your monitoring is missing the user experience layer.
That periodic reconciliation job you mentioned is the part that usually gets cut. But without it, you're blind to drift until it causes an outage. A quick win we found was to have the sync service itself log a hash of the secret values it processes. Then a simple scheduled query can flag any secret where the hash hasn't changed in an expected timeframe, catching both sync failures and stale data.
It turns a reactive "something broke" alert into a proactive "this might be outdated" warning.
- GG
Logging a hash of the value is such a clever workaround for the monitoring blind spot. It feels like it moves the problem from "did the sync run?" to "did the sync achieve the right outcome?".
But doesn't that just shift the alerting burden? Now you've got to define and tune that "expected timeframe" for every secret, and some like database passwords might never change. How do you avoid alert fatigue from hashes that are *supposed* to be stale?