Having just completed a multi-year phased migration of a 500+ service estate to AWS, one of the most persistent operational burdens post-migration was not the infrastructure itself, but the compliance and security posture enforcement. We deployed Banyan Security for Zero Trust access, and while its policy engine is robust, the reality of day-to-day operations involved a constant stream of tickets from developers whose pods or VMs were flagged for non-compliance (e.g., missing agents, outdated vulnerability scans, missing security tags). My team was essentially acting as a manual remediation layer, which was neither scalable nor a good use of senior architect time.
The core issue was the feedback loop. A Banyan TrustScore would drop, access would be restricted per our policies, and a Jira ticket was auto-generated. However, the remediation steps were often repetitive and could be codified. We decided to build an automated remediation orchestrator that sits between Banyan's alerts and our ticketing system. The result has been an 80% reduction in related tickets, as the majority of common failures are now fixed before a human needs to look at them.
Our system is built on a few key AWS services and leverages Banyan's API. The workflow is as follows:
1. **Event Capture:** Banyan webhooks are configured to send `trustscore.changed` events to an AWS EventBridge event bus.
2. **Evaluation & Routing:** An EventBridge rule filters for events where the `new_trustscore` is below our threshold (e.g., 7). It routes the event payload to a Step Functions state machine for orchestration.
3. **Remediation Logic:** The Step Functions state machine is the core. It uses Lambda functions to:
* Fetch detailed device/service posture details from Banyan's `GET /v1/device_posture` or `/v1/service_posture` APIs.
* Parse the `failure_reasons` array.
* Execute specific remediation actions based on the failure reason.
Here is a simplified example of our Step Functions definition (in ASL) showing the branch logic:
```json
{
"Comment": "Remediate Banyan Posture Failure",
"StartAt": "GetPostureDetails",
"States": {
"GetPostureDetails": {
"Type": "Task",
"Resource": "arn:aws:states:::lambda:invoke",
"Parameters": {
"FunctionName": "GetPostureDetailsFunction",
"Payload": {
"device_id.$": "$.detail.device_id"
}
},
"Next": "AnalyzeFailures"
},
"AnalyzeFailures": {
"Type": "Choice",
"Choices": [
{
"Variable": "$.failure_reasons",
"StringMatches": "*Banyan Agent not heartbeating*",
"Next": "RemediateAgent"
},
{
"Variable": "$.failure_reasons",
"StringMatches": "*Vulnerability scan older than*",
"Next": "TriggerVulnScan"
},
{
"Variable": "$.failure_reasons",
"StringMatches": "*Missing security tag*",
"Next": "ApplyTag"
}
],
"Default": "CreateManualTicket"
},
"RemediateAgent": {
"Type": "Task",
"Resource": "arn:aws:states:::lambda:invoke",
"Parameters": {
"FunctionName": "RemediateAgentFunction",
"Payload": {
"instance_id.$": "$.detail.cloud_instance_id",
"device_id.$": "$.detail.device_id"
}
},
"Next": "VerifyRemediation"
},
"CreateManualTicket": {
"Type": "Task",
"Resource": "arn:aws:states:::lambda:invoke",
"Parameters": {
"FunctionName": "CreateJiraTicketFunction"
},
"End": true
},
"VerifyRemediation": {
"Type": "Wait",
"Seconds": 120,
"Next": "RecheckPosture"
},
"RecheckPosture": {
"Type": "Task",
"Resource": "arn:aws:states:::lambda:invoke",
"Parameters": {
"FunctionName": "GetPostureDetailsFunction"
},
"Next": "EvaluateResult"
},
"EvaluateResult": {
"Type": "Choice",
"Choices": [
{
"Variable": "$.trustscore",
"NumericGreaterThanEquals": 7,
"Next": "Success"
},
{
"Variable": "$.trustscore",
"NumericLessThan": 7,
"Next": "CreateManualTicket"
}
]
},
"Success": {
"Type": "Succeed"
}
}
}
```
The Lambda functions perform actions like:
* **`RemediateAgentFunction`:** For EC2 instances, it uses SSM Run Command to restart the Banyan service. For EKS pods, it annotates the deployment to force a rolling restart.
* **`TriggerVulnScanFunction`:** Invokes a separate vulnerability scanning pipeline via Step Functions, passing the asset identifier.
* **`ApplyTagFunction`:** Uses the AWS Resource Groups Tagging API or Kubernetes API to apply the required tags, based on the asset type.
Key architectural considerations and pitfalls we encountered:
* **Idempotency is Critical:** Every remediation Lambda must be idempotent. We use a combination of the Banyan `device_id` and a failure reason code as a deduplication key, storing state in a short-lived DynamoDB table to prevent loops.
* **Security Context:** The Lambda functions need permissions for both Banyan (API key stored in Secrets Manager) and the target infrastructure (EC2, EKS, etc.). This requires careful, scoped IAM roles and pod identities.
* **Verification Wait Time:** The `Wait` state between remediation and re-check is crucial. Banyan's TrustScore updates are not instantaneous; we found 90-120 seconds to be reliable.
* **Escalation Path:** The `CreateManualTicket` state is the essential safety net. Any failure pattern the system doesn't recognize, or a remediation that fails twice, automatically kicks out a well-formatted Jira ticket for the security team.
The cost is negligible—perhaps a few dollars a month for Lambda and Step Functions usage. The ROI, however, has been substantial. Developer productivity improved as access blocks were shorter, and my team's cognitive load decreased significantly. This pattern has proven so effective we're now exploring its application to other areas of our security toolchain. The principle is sound: use the observability provided by your Zero Trust platform not just for alerting, but to drive a closed-loop, automated remediation system.
That's a fantastic outcome. I'm curious, how did you handle the safety measures for your orchestrator? You mentioned it sits between alerts and the ticketing system.
When we built something similar for our alert pipeline, we had to implement a circuit breaker pattern. We tracked the number of consecutive automatic remediation attempts per resource, and if it failed more than twice, it would stop trying and open a ticket. This prevented runaway automation from hammering a resource that was genuinely broken. Did you run into anything like that?
Sleep is for the weak
An 80% reduction sounds great, but you didn't mention what happens when your orchestrator makes the wrong fix. Automating a remediation based on an alert assumes the alert is always right.
Banyan flags something for a missing tag, your script applies a tag, but what if the flag was because the resource shouldn't exist at all? You just automated the creation of technical debt.
your mileage will vary
You're absolutely right to flag that. Automating remediation based on a failed check is only as smart as the check itself. In our case, the orchestrator only acts on a very narrow set of well understood failures, like a missing agent that we know how to reinstall cleanly.
For something like a missing tag, you need a layer of context. Our script won't apply a tag unless the resource passes a separate validation check confirming it's a sanctioned, active service. It prevents exactly that scenario of automating technical debt. The key is not just reading the alert, but having a decision logic that asks "should this resource even be here?" before you touch it.
Keep it civil, keep it real
Exactly. That separate validation check is the core of a safe automation loop.
You can't just trust the first alarm. We do the same with our pipeline's security scans. A vulnerability flag doesn't trigger a fix. It triggers an inventory lookup to see if the flagged component is even in a deployed service. If it's not, we mute the alert. If it is, then we run the fix.
> "should this resource even be here?"
We call that a context gate. Without it, you're just building a faster shovel for your own grave.
That context gate is the part everyone wants to skip because it's the boring, hard work. You have to build and maintain that inventory lookup. I've seen projects die because they didn't budget for that. The automation script is sexy, but the asset database is the real project.
"Faster shovel for your own grave" is a perfect description for most half-baked automation. How do you handle drift between your context gate source and reality? Our CMDB is wrong about 20% of the time.
trust but verify
The "context gate" idea makes a lot of sense. I've seen similar logic work for our team's time tracking alerts, where a missing entry triggers a check against the project calendar before sending a nudge.
How do you stop the inventory lookup itself from becoming a bottleneck?
That initial setup with auto-generated tickets but manual remediation is exactly what eats up cycles. Reducing it by 80% is huge.
Your point about codifying repetitive steps is key. I'm curious, when you say the orchestrator sits between alerts and the ticketing system, what's your logic for when it *does* still create a ticket? Is it only for failures it doesn't recognize, or do you have a manual review flag for certain resource types?
Benchmarking my way to better decisions
> an 80% reduction in related tickets
That's the vendor-slide metric that always gets trotted out. Sure, your *ticket* volume dropped, but did your actual *problem* volume drop by the same amount?
I'm betting your script is just automating fixes for the low-hanging, repetitive alerts. The other 20% of tickets you're still getting are now probably the complex, time-sucking ones that take ten times longer to diagnose. You might have actually *increased* your team's cognitive load while improving a KPI.
Also, you stopped mid-sentence: "built on a few key A..." Am I supposed to guess? Lambda? Ansible? I'm assuming the "A" is for "Assumptions" about your environment staying perfectly static.
Trust but verify.