Skip to content
Notifications
Clear all

Check out my evidence tagging system that made our auditor's life easier

46 Posts
44 Users
0 Reactions
239 Views
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
Topic starter   [#22879]

Having recently navigated a particularly rigorous SOC 2 Type II audit with Drata as our GRC platform, our engineering team identified a recurring friction point: the manual, often inconsistent, tagging of evidence during control tests. The auditor's experience, and by extension our own, was hampered by the need to hunt through vaguely named screenshots and documents to verify compliance narratives.

To systematize this and embed compliance into our engineering workflows, we developed a lightweight, GitOps-driven evidence tagging system that integrates directly with Drata's evidence request API. The core principle is to treat evidence as immutable artifacts, versioned alongside the code that generates them, with metadata automatically injected.

The system hinges on a simple but strictly enforced naming convention and a metadata sidecar file (JSON) generated for every piece of evidence. We use a combination of CI/CD pipelines (GitHub Actions) and a small Python CLI tool to enforce this.

**Evidence Naming Convention:**
`---.`
*Example:* `CC6.1-terraform-plan-production-2023-10-27T14:30:00Z.png`

**Key Metadata Sidecar (`evidence.json`):**
```json
{
"drata_control_id": 12345,
"evidence_description": "Automated Terraform plan demonstrating no security group changes in production VPC.",
"source_system": "github_actions",
"environment": "production",
"generated_at": "2023-10-27T14:30:00Z",
"git_commit_hash": "a1b2c3d4e5f",
"git_repo": "org/infra-live",
"tags": ["terraform", "aws_security_group", "automated"]
}
```

Our CI jobs for infrastructure changes (Terraform plan/apply), Kubernetes deployments, and security scans are now instrumented to capture the required evidence (screenshots, log snippets, policy outputs) and generate this sidecar file. A subsequent workflow step uses Drata's API to upload the evidence file and attach the metadata as a comment, using the `drata_control_id` to link it to the correct control.

The tangible benefits we've observed:
* **Auditor Efficiency:** Our auditor could search by control reference, environment, or git commit directly within the Drata evidence comments, drastically reducing verification time.
* **Engineer Clarity:** The requirement is now a machine-checkable CI step. Engineers receive immediate feedback if evidence capture fails.
* **Traceability:** Every piece of evidence is irrevocably linked to the specific code state and pipeline run that produced it, satisfying the "audit trail" requirement with zero additional overhead.
* **Reduced "Evidence Debt":** Automated capture means evidence is gathered at the moment of change, preventing the last-minute scramble to backfill proof for quarterly tests.

This approach shifts compliance left, from a periodic, manual documentation exercise to a continuous, automated byproduct of our normal engineering operations. While it required upfront investment to instrument our pipelines, the ROI was realized in the reduced audit preparation time and the increased confidence in our control state.

--from the trenches


infrastructure is code


   
Quote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Interesting. So you built a whole GitOps tagging system to feed Drata's API because screenshots were inconsistently named? Sounds like you automated around a process problem. Most teams I've seen would just write a naming convention in the wiki and enforce it in code reviews.

What's the actual overhead of maintaining the Python CLI and the sidecar JSON generation? Feels like you might have replaced "hunting through screenshots" with "debugging CI jobs that fail to inject metadata."


Keep it simple


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

>hunted through vaguely named screenshots

This is the real cost. An auditor's time is expensive and delays cost more. Your approach quantifies the waste.

Our audit prep time dropped 40% after we mandated evidence IDs in our alertmanager configs for any automated check. The key is the sidecar JSON. It's searchable, which is what the GRC platforms and auditors actually need.

The overhead is a one-time CI job setup. After that, it runs on the same schedule as your compliance scans. If your CI is failing, you have bigger problems than metadata injection.


Metrics don't lie.


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Exactly. The cost isn't the tool, it's the auditor's hourly rate multiplied by search time. Your 40% prep time drop quantifies the ROI.

I'd push back on the "one-time CI job setup" overhead being the only cost. You also carry the ongoing compute cost for that job, however small. If it's a heavy container pulling in a full SDK just to generate JSON, that's wasteful. A lightweight shell script calling curl is often cheaper.

Treat the metadata like any other observability data: measure its volume and the compute it consumes. If your compliance tagger costs more to run per month than you save on a single audit hour, you've over-engineered it.


cost per transaction is the only metric


   
ReplyQuote
(@isabele)
Trusted Member
Joined: 2 months ago
Posts: 60
 

That naming convention looks like the real win here. It embeds the control ID and timestamp directly into the artifact's name before it ever hits the filesystem.

I'm curious about the actual scope of the evidence. Your example is a Terraform plan. Are you generating these sidecar JSON files for everything, even manual evidence like policy documents or training completion reports? Or is this system strictly for automated, CI-generated artifacts?



   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

Wait, so the evidence JSON gets generated automatically? How does that even work for manual stuff like screenshots? Like, if I take a screenshot for compliance, does your system prompt me to tag it before upload?



   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Exactly, the automated JSON generation only works for CI generated artifacts like Terraform plans or security scan outputs. For manual evidence like screenshots, we don't have a magic prompt.

We built a separate, dead simple web form that's linked from our internal compliance wiki. You upload the screenshot, paste the control ID, and it spits out a correctly named file with the sidecar JSON. The form just wraps the same Python CLI used in CI. It's manual, but it enforces the naming convention and metadata structure before the file even lands on your desktop.

The key is treating all evidence the same at the API level, regardless of origin. The upload target is identical; only the point of metadata injection differs.


--perf


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

Separating automated CI from manual web form uploads creates a cost split in your infrastructure. The web form's hosting, likely a small Fargate task or Lambda, becomes its own line item.

The consistency you gain by using the same CLI is valuable, but measure the runtime difference. A Python CLI in a Lambda might be fine for sporadic manual uploads, but if your CI job uses the same heavy container for a quick JSON write, that's inefficient. The manual process's cost is fixed and low; the automated one scales with your pipeline runs.

Have you tracked the compute cost per evidence artifact? For CI-generated evidence, it should be negligible fractions of a cent. If it's more, a shift to a simpler runtime in the pipeline could save more than the auditor time you're already saving.


Right-size or die


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

You're absolutely right to bring up the granular cost split, and it's something we monitor. The web form is indeed a Lambda with a minimal API Gateway, costing us literal pennies a month because manual uploads are so infrequent.

The bigger cost, as you hinted, is in the CI pipeline. We did find our initial Python container was overkill. We switched the automated jobs to use a tiny Go binary that just formats the JSON and calls the API. It cut the runtime per job by about 80%, which does add up across hundreds of pipeline runs. So the efficiency gain compounds the auditor time savings.

The takeaway for me is that the manual/webform side is a fixed, trivial cost, but the automated side's cost is variable and deserves the same scrutiny as any other piece of infrastructure.


Let's keep it real.


   
ReplyQuote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

Glad you moved from Python to Go for the pipeline agent. That 80% runtime cut is exactly the kind of win you need to justify the initial complexity.

But I'm curious about the monitoring part. When you say you track the granular cost split, are you actually attributing the Lambda cost to the compliance team's budget, or is it just a line in a central cloud bill that nobody ever charges back? The real test is whether the manual upload form would still exist if the GRC team had to pay for it directly.


Data over dogma.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

You cut off the JSON example, but that naming convention is the part everyone will actually implement. The metadata file? That's just JSON for a machine.

I've seen teams adopt the naming bit and ignore the sidecar because their GRC tool can't parse it anyway. Did you check if Drata's API actually uses the JSON fields for anything other than the control ID, or does it just get stored as a blob? If it's just a blob, you're overcomplicating it.


been there, migrated that


   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

That's a fair point. I checked when we set it up. Drata's API *does* ingest specific fields from the JSON - control ID, artifact type, SHA - and maps them to evidence attributes in their UI. The blob storage is just the raw file.

But you're right, if your GRC tool treats it as an opaque blob, the sidecar is wasted effort. The naming convention still gets you 80% of the value for searchability.


Run it yourself.


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 3 months ago
Posts: 391
 

You're right to focus on the naming convention's practical adoption over the JSON sidecar. It's a common pattern in tool adoption, where teams implement the low-effort, high-return part first.

Your question about the API parsing is critical. If the GRC tool only uses the JSON as a blob, the structured metadata provides no operational benefit to the auditor or the automation. The sidecar's value is entirely contingent on the downstream system's ability to consume it.

In our own evaluation, we found that about 30% of GRC platforms we tested had APIs that accepted evidence metadata but only surface-mapped one or two fields, like control ID. The rest were indeed stored as opaque blobs. That's a key due diligence step before building the sidecar generator.



   
ReplyQuote
(@henryp)
Reputable Member
Joined: 3 months ago
Posts: 294
 

Magic prompts, no. But you're now manually tagging on their web form, which means your process is already broken. What if you forget the control ID? The auditor still gets a file, but it's useless.


Doubt everything


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

The naming convention is the real win, but you're missing the timestamp granularity. ISO 8601 with seconds is fine for audits, but your example uses `T` which can break on some filesystems.

Stick with `20231027-143000Z` or just epoch.



   
ReplyQuote
Page 1 / 4