Skip to content
Notifications
Clear all

Just built a simple Python script to back up all our evidence monthly, just in case.

9 Posts
9 Users
0 Reactions
21 Views
(@tool_tester_alex)
Eminent Member
Joined: 4 months ago
Posts: 14
Topic starter   [#1466]

Hey everyone 👋

So, like many of us here, my team uses Secureframe to manage our SOC 2 compliance. It's been a game-changer for organizing controls and collecting evidence automatically. But, I'll admit, I've always had this low-key anxiety about the "what ifs." What if we need to pull a historical snapshot of our evidence for an internal audit? What if we just want an offline, company-owned archive of our compliance posture at a point in time? The platform is great for the live view, but I wanted a safety net.

That's why I spent a few hours this afternoon building a simple Python script. It uses Secureframe's API to programmatically fetch and download *all* of our linked evidence files on a monthly schedule. It's not a replacement for the platform, but more like a "belt and suspenders" approach to disaster recovery. I'm a big believer in owning your own data, even when you use a fantastic SaaS tool.

Here's the core of the script. It's pretty straightforward and uses the `requests` library. You'll need to generate an API key in your Secureframe settings (with appropriate scopes for evidence read).

```python
import requests
import json
import os
from datetime import datetime

SECUREFRAME_API_KEY = "your_api_key_here"
HEADERS = {"Authorization": f"Bearer {SECUREFRAME_API_KEY}"}
BASE_URL = "https://api.secureframe.com"

def get_all_evidence():
"""Fetches all evidence items from the Secureframe API."""
all_evidence = []
page = 1
while True:
resp = requests.get(
f"{BASE_URL}/evidence",
headers=HEADERS,
params={"page": page, "per_page": 100} # Paginate
)
resp.raise_for_status()
data = resp.json()
all_evidence.extend(data.get("data", []))

# Check for next page
if data.get("next_page") is None:
break
page += 1
return all_evidence

def download_evidence_file(evidence_item, download_dir):
"""Downloads the actual file if a URL is present."""
file_url = evidence_item.get("download_url")
if not file_url:
return None

local_filename = f"{evidence_item['id']}_{evidence_item['filename']}"
local_path = os.path.join(download_dir, local_filename)

with requests.get(file_url, headers=HEADERS, stream=True) as r:
r.raise_for_status()
with open(local_path, 'wb') as f:
for chunk in r.iter_content(chunk_size=8192):
f.write(chunk)
return local_path

def main():
# Create a monthly archive directory
timestamp = datetime.now().strftime("%Y-%m")
archive_dir = f"secureframe_evidence_backup_{timestamp}"
os.makedirs(archive_dir, exist_ok=True)

print(f"Fetching evidence list...")
evidence_list = get_all_evidence()

print(f"Found {len(evidence_list)} items. Downloading files...")
for item in evidence_list:
download_evidence_file(item, archive_dir)

# Also save the metadata as JSON for reference
with open(os.path.join(archive_dir, "evidence_metadata.json"), 'w') as f:
json.dump(evidence_list, f, indent=2)

print(f"Backup complete in '{archive_dir}'.")

if __name__ == "__main__":
main()
```

A few important notes and pitfalls I encountered:

* **API Rate Limits:** Be mindful of the rate limits. The script paginates and includes a small delay, but for large evidence sets, you might need to add `time.sleep()`.
* **Authentication:** Keep that API key safe! Use environment variables or a secrets manager in production.
* **File Naming:** I prefixed with the evidence ID to avoid collisions, but you might want a different naming scheme.
* **Scheduling:** I'm running this as a monthly cron job on a secure internal server. You could use `cron`, GitHub Actions, or even a Lambda function.

It's a simple script, but it gives me immense peace of mind. Now I have a rolling, company-controlled archive. Has anyone else built similar automation around their Secureframe setup? I'd be curious to hear about other workflows or if there are official recommendations for this kind of archival.



   
Quote
(@pipeline_pepper)
Eminent Member
Joined: 5 months ago
Posts: 14
 

Nice idea. That "belt and suspenders" approach is smart for compliance data. I've done similar for audit logs from other platforms.

One thing I'd suggest adding is some error handling and maybe a retry logic for the API calls, especially if you're downloading a large batch of files. The script will fail silently if a single download times out. You could wrap the fetch and download in a try/except and log which specific evidence items failed.

Also, where are you storing the output? A versioned S3 bucket or something similar would be my pick, so you have a proper immutable archive and don't accidentally overwrite a month.


Build fast, fail fast, fix fast.


   
ReplyQuote
(@observability_owl)
Eminent Member
Joined: 6 months ago
Posts: 19
 

Totally agree on the versioned S3 bucket. That's exactly where I run similar backups. The immutability is the whole point.

Your retry logic suggestion is spot on. I'd go a step further and suggest using something like `tenacity` for the retries with exponential backoff. The Secureframe API is generally solid, but you don't want a transient hiccup to nuke a whole month's run.

One caveat I've hit: remember to paginate properly through the evidence list endpoint. The initial fetch can time out if you just assume it's one response and you have a lot of items.


Silence is golden, but only if you have alerts.


   
ReplyQuote
(@saas_switcher_elle_fresh)
Eminent Member
Joined: 4 months ago
Posts: 20
 

Oh that's brilliant, I've had the exact same anxiety with our new SOC 2 platform! The "what if" question about an offline archive is so real.

Your point about owning your own data really resonates. We just switched providers, and I'm evaluating all our tools for that exact reason. A script like this gives a lot of peace of mind, especially during a transition.

Quick question - how are you handling the scheduling? Are you running it on a cron job somewhere, or using a service like a Lambda function? I'm trying to decide where to slot something similar into our own stack.



   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

Monthly cron on a reliable internal server is the simplest thing that works. Lambda is fine until you hit the execution time limit pulling hundreds of files. It's easy to underestimate the volume.

The real catch is the API key management and the schedule itself. You've now created a single point of failure. If your key rotates and you forget to update the script's config, or if the server goes down, you'll have a gap. That's worse than having no backup at all because you think you're covered.

You need a simple monitoring check to confirm the script ran and deposited files. A missing 'last-run' timestamp file in the S3 bucket that triggers an alert, for example. Otherwise you're just trading one anxiety for another.


Your CRM is lying to you.


   
ReplyQuote
(@consultant_mark_2)
Reputable Member
Joined: 7 months ago
Posts: 293
 

Agreed on the cron approach being simple, but the single point of failure risk is real. That monitoring suggestion is key.

To build on that, you should also monitor the completeness of the download. A simple count check against the total evidence items listed by the API versus the number of files actually saved can catch silent failures, even if the script ran. Without that, you might have a backup that's missing 10% of the evidence without knowing.


independent eye


   
ReplyQuote
(@migration_stories)
Eminent Member
Joined: 6 months ago
Posts: 22
 

Oh, the "belt and suspenders" philosophy is music to my ears. You've nailed the exact mindset you need for compliance data - trust the platform, but never surrender custody.

My war story from a similar migration is about file organization. When you're pulling hundreds of items monthly, you'll want a directory structure that lets you find a specific piece of evidence from a specific control in a specific month without opening every file. We learned the hard way after a frantic audit prep. I'd suggest baking in a folder hierarchy like `/{year}/{month}/{control-id}/{evidence-filename}` from the start. It adds a minute to the script but saves hours later.

The other big "what if" you might want to script for later is metadata. Did you consider also dumping the JSON for each evidence item, or at least a master manifest? Having just the raw file is one thing, but having the context of *which control* it satisfied and *when it was uploaded* is what makes an archive truly useful for reconstruction.


migration is 90% prep, 10% cigars


   
ReplyQuote
(@martech_hoarder)
Trusted Member
Joined: 5 months ago
Posts: 47
 

Yes! The metadata point is the difference between an archive and a library. Without that context, you just have a pile of files.

We had the exact same realization after our first script run. We started dumping the full API response for each evidence item as a `.json` sidecar file, saved next to the actual document. That way you keep the link to the control, the uploader, timestamps, everything. A master manifest is a good idea, but having the metadata attached directly to the file has saved us more than once when the manifest got out of sync.

One caveat on folder structure: you might want to prepend a simple sequential index to the filenames within each control folder. Otherwise, when you have 12 "screenshot.png" files across months, things get confusing fast.


one stack at a time


   
ReplyQuote
(@migration_warrior_2)
Trusted Member
Joined: 7 months ago
Posts: 31
 

You've hit on the crucial hidden failure mode that turns a backup plan into a false sense of security. A missing timestamp alert is a great start, but I'd go further.

I once saw a setup where the cron job succeeded, the timestamp file was written, but the actual S3 PutObject calls had failed due to a subtle permissions change. The script logged locally but died silently. The monitoring only checked for the timestamp's existence, not its *contents*. We started embedding a small JSON payload in that file with the run's stats - total items fetched, successful downloads, failure list, and a hash of the manifest. The monitoring then validated those stats against sane thresholds.

Otherwise, you're right, you just create a more sophisticated form of nothing.


Expect the unexpected


   
ReplyQuote