Skip to content
Notifications
Clear all

Just built a custom control mapping for SOC 2 using their API - here's the script.

31 Posts
31 Users
0 Reactions
75 Views
(@clairen)
Reputable Member
Joined: 3 months ago
Posts: 390
Topic starter   [#27296]

I've been knee-deep in our SOC 2 prep, and one of the more tedious parts has always been mapping our internal security controls to the official Trust Services Criteria. The out-of-the-box mappings in Secureframe are good, but our setup is... particular. We have a lot of custom, in-house tooling for deployment and access reviews.

Since I live in the terminal, I decided to use Secureframe's API to automate it. The goal was to create a script that takes our internal control IDs and links them to the relevant Secureframe control objects, saving our compliance team hours of manual clicking.

Here's the gist of the approach:
* The API is RESTful and well-documented. The key endpoint for this is `POST /v2/control_mappings`.
* You need to first fetch the list of existing Secureframe controls (via `GET /v2/controls`) to get their UUIDs.
* The script then reads a CSV where each row has our internal control code and the corresponding Secureframe control ID(s). It's basically a data pipeline: extract from CSV, transform/lookup IDs, load to API.

The main gotcha was handling the many-to-many relationships. One of our internal controls might map to three SOC 2 criteria, and the API payload needs to reflect that. Also, error handling is crucial—you don't want to silently fail on a partial batch.

This has been a game-changer for our workflow. It ensures consistency and lets us version our mappings in git. Curious if anyone else has tried something similar? I'm wondering about the long-term maintenance, especially when Secureframe updates their standard library. Do you just re-run the mapping script and review the diffs?

—Claire



   
Quote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

API's well-documented? You must have gotten a different one. The rate limiting on that GET /v2/controls endpoint is punitive. You'll fetch fifty controls and then get throttled for an hour.

Did you handle pagination? Their default limit is low and you'll miss half your controls if you just call it once.

And CSV? That's asking for trouble with special characters in the control descriptions. A real script would use something that can handle actual data escaping, or at least validate the fields before POST.


-- old school


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

The many-to-many mapping is the critical piece. I've found you also need to plan for the historical record. If you ever need to *update* a mapping, you're not just patching it. Many audit logging systems, including some SIEM integrations, treat a changed control mapping as a deactivation of the old link and a creation of a new one. Does the Secureframe API give you timestamps on the mapping object itself? You'd want to capture those for your own audit trail.


Logs don't lie.


   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

That's an excellent point about the audit trail, and it's one many teams overlook until they're in a remediation meeting. The Secureframe API does include `created_at` and `updated_at` timestamps on the mapping object. However, if the API performs a soft delete on the old mapping when you update, you might lose visibility unless you're logging the full API request and response payloads separately.

You'd need to design your script to capture the mapping object's entire state, including those timestamps, and push it to your own immutable log before making any update call. Treating it as a deactivation and new creation is the right mental model for compliance purposes, even if the vendor's API abstracts it behind a single PUT endpoint.


—at


   
ReplyQuote
(@infra_architect_6)
Reputable Member
Joined: 5 months ago
Posts: 259
 

Mapping many-to-many relationships is the core complexity. The API payload structure for `POST /v2/control_mappings` likely expects an array of Secureframe control UUIDs for each internal control. Your script's transformation layer must be idempotent to handle partial failures. If the CSV has an internal control listed across multiple rows, you'll need to consolidate those Secureframe IDs into a single array before the POST, otherwise you'll create duplicate mapping entries.

Also, consider how this mapping pipeline integrates with your existing infrastructure. If your internal control IDs originate from a Git repository or a CMDB, the script should pull directly from that source of truth, not an intermediate CSV. This avoids data drift. You could structure it as a Terraform provider or a Kubernetes controller that reconciles the desired state with the Secureframe API, applying GitOps principles to compliance mapping.

The payload size could become an issue. If a single internal control maps to dozens of Secureframe controls, you might hit undocumented limits on the array length. You'd need to implement batching or chunking within the script's load phase.



   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

You cut off at the API payload, but that's the critical part. You need to show the exact structure of your JSON payload for `POST /v2/control_mappings`. A malformed array there is the difference between a clean run and a data integrity mess.

Did you validate that the mapping was actually created successfully? The API might return a 200 but with errors nested in the response body. Always log the full JSON response for each call.


Show me the query.


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Right? That payload structure is everything. For ours, the POST body was a list of objects, each needing the internal control `id` (our own identifier) and a `secureframe_control_uuids` array. The key was making sure that array was never empty and each UUID was valid from their controls endpoint.

We absolutely logged every single response, even on 200. Found a few "successful" calls that had warnings about deprecated control IDs being referenced. Those would've been silent failures. We piped the logs to a separate SIEM channel just for compliance automation events.

And for validation, we added a step that fetches the mapping list after the script runs, comparing the count. If the numbers didn't match, it'd alert and halt. Saved us from a mess more than once when the API had a hiccup mid-run.


Pipeline is king.


   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 3 months ago
Posts: 295
 

"The API is RESTful and well-documented." That's a charitable reading. I suppose documentation is a separate skill from building an API that's actually pleasant to use.

Saving your compliance team hours of manual clicking is a noble goal, but I hope you've factored in the hours you'll spend maintaining this script when their API changes, or when your CSV format inevitably diverges from the source of truth.


Beware of free tiers


   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

Saving your compliance team hours, you say. I'll be the cynic here: have you actually tracked that saved time, or is this another case of "automation for automation's sake" that just moves the manual effort from the compliance team to the platform team? The script isn't free. Someone's now on the hook for the API changes, the CSV drift, and the inevitable "why did the mapping break?" tickets.

I'd be more convinced if you showed the billing data - like the reduction in Secureframe professional services hours or the actual FTE time reallocated. Otherwise, this is just shifting costs, not cutting them.


cost_observer_42


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

You're right about the many-to-many relationships being the core architectural challenge. The script's transformation logic has to be absolutely idempotent, especially when ingesting from a CSV. A row-based processing model will create duplicate mapping entries unless you aggregate all `secureframe_control_uuids` by your internal `id` before making the API call.

Building this as a standalone script that parses a CSV is a good start, but the real maintenance cost, as others have hinted, comes from that CSV becoming an out-of-sync artifact. The next logical step is to source your internal control IDs directly from your policy-as-code repository or CMDB via their API, making the script a stateless integration layer between two declarative systems. That eliminates the data drift problem entirely.


infra nerd, cost hawk


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Oh, I'm totally with you on tracking the actual savings. We learned that the hard way on our last SOC 2 audit cycle.

We actually did measure it, but in a roundabout way. Our compliance lead tracks time spent in Jira labels. Before automation, the "control mapping" label had about 40-50 hours per audit prep cycle, mostly manual clicking and validation. After we built a pipeline from our policy repo to Secureframe, that dropped to under 5 hours for a cycle, and those hours were mostly reviewing the automated diffs. The cost shifted to my team, sure, but it was a net reduction because my team's 10 hours of script maintenance was less than the 35+ hours saved on their side.

The real win, though, was in accuracy. Manual mapping had a consistent error rate that caused rework. The automated version eliminated that entirely, and *that's* where the hidden cost was buried. You can't always see that in a billing line item.


Backup first.


   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

That CSV-to-API pipeline is exactly how we started too! Getting the UUIDs from `/v2/controls` first is crucial, but I'd recommend caching that list locally for the script's run. The `GET` call is cheap, but if you're processing a huge CSV and doing a lookup for every single row, you could hit rate limits unintentionally. A simple JSON file dump at the start saves a lot of potential headache.

And totally feel you on the many-to-many complexity. When you say "the API payload...," I'm guessing you structured it as an array of objects? The big lesson for us was that we had to deduplicate aggressively on our internal ID *before* building that final payload. Our first run created dozens of duplicate mappings because the same internal control appeared across multiple CSV rows for different Secureframe criteria. A quick aggregation step in the transform phase fixed it.


null


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

Yep, the local cache is a lifesaver. I actually take it a step further and version that JSON dump alongside our script in git. That way, if the Secureframe API is having a bad day during an urgent run, we can still execute with a known-good snapshot from last week, then update the cache later. It's a nice fallback.

And on deduplication, your point about aggregation is spot on. We made the same mistake early on. Our script now uses a simple dictionary keyed by the internal ID to accumulate UUIDs as it reads the CSV, only making the final POST call after the entire file is processed and deduped. It's the only safe way with a row-based source.


api first


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

That versioning trick is clever. Have you had a scenario where the API schema changed but your cached data still worked? I'd worry about using an old snapshot if new required fields were added.



   
ReplyQuote
(@briang)
Estimable Member
Joined: 3 months ago
Posts: 119
 

Good question. We haven't seen a schema change break it yet, because our script's payload only uses the control UUIDs from that cached list. As long as the IDs themselves are still valid, the mapping call seems to work.

But your worry is valid. If they ever changed the required structure of the POST body itself, the old cache wouldn't help there. It just prevents lookup failures during that one step.

How do you usually monitor for API changes? Just release notes?



   
ReplyQuote
Page 1 / 3