Skip to content
Notifications
Clear all

Step-by-step: Handling an exception or failed test without panicking the team

2 Posts
2 Users
0 Reactions
28 Views
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
Topic starter   [#12473]

A failed control test in Drata—or any compliance automation platform—often triggers an alarmist response, especially from teams new to continuous compliance. The immediate instinct is to treat it as a security incident, flooding Slack channels and summoning emergency meetings. This is counterproductive. In a mature deployment, exceptions and failures are not emergencies; they are data points in a feedback loop for your governance posture.

The core principle is to decouple the *alert* from the *response*. A failed test indicates a deviation from your documented policy or a technical misconfiguration. Your first step is not to "fix" it blindly, but to classify it. Here is a structured workflow I've implemented for clients to triage Drata test failures without inciting panic:

1. **Immediate Triage (Owner: Compliance Lead or On-Call Engineer)**
* Access the Drata dashboard and navigate to the specific failed test.
* Determine the failure category:
* **False Positive:** The test is incorrect. Example: A test checking for disk encryption on a stateless, ephemeral Kubernetes pod.
* **Expected Exception:** The deviation is authorized and documented. Example: A non-compliant legacy system scheduled for decommissioning.
* **True Positive:** A genuine policy violation or configuration drift.
* This classification must happen within the first business hour, but not via all-hands communication. Use a dedicated, low-volume compliance channel.

2. **Containment & Documentation**
* For **False Positives**, update the test logic or its scope in Drata. This often involves refining resource queries.
```yaml
# Example: Adjusting a Drata test query for AWS to exclude specific S3 buckets used for transient logging.
# Original might scan all S3 buckets. Refined query adds an exclusion filter.
resources:
- type: aws_s3_bucket
filters:
- not:
- tag:Compliance-Exclusion: "transient-logs"
```
* For **Expected Exceptions**, you must have a pre-defined process. Immediately create an Exception Record in your GRC tool (or even a tracked issue in Jira). Link the Drata failure to this record. The record should contain the business justification, risk acceptance, owner, and expiration date. Then, in Drata, you can add a note to the test and, if appropriate, temporarily suppress it for that specific resource—*never* disable the test entirely.
* For **True Positives**, initiate your standard incident or change management process. The key is that this is now a tracked remediation item, not a shouting match.

3. **Communication Protocol**
* **Status Page:** Maintain an internal dashboard (can be a simple, automated wiki page) that reflects the health of your compliance controls. Green/Amber/Red status should be based on *unresolved true positives*, not total failures.
* **Reporting:** Daily digest emails to the security and engineering leadership summarizing new failures, their categories, and resolution ETA. This replaces ad-hoc, panic-driven messages.
* **Blameless Review:** Weekly, review the root cause of true positives. Was it a deployment flaw, a policy gap, or a training issue? This turns failures into systemic improvements.

The goal is to engineer the panic out of the system. By treating Drata as a source of telemetry for your security posture—akin to application logs or performance metrics—you normalize failures as part of the operational lifecycle. This requires upfront work to define the triage workflow and exception process, but it pays dividends in team morale and audit readiness. The auditor isn't looking for perfection; they're looking for evidence of a controlled, managed process. A well-documented exception is often more compelling than a hidden, silently "fixed" failure.

- Mike


Mike


   
Quote
(@crusty_pipeline_v2)
Reputable Member
Joined: 5 months ago
Posts: 338
 

Good start. Your third category is missing - the "legitimate failure" that requires a real fix. Teams often stall because they can't decide between "expected exception" and "actual problem."

For false positives, I add a rule: if it's a platform bug, log a ticket with the vendor *and* immediately create a custom monitor to catch the real condition. Don't wait for their fix.


slow pipelines make me cranky


   
ReplyQuote