Skip to content
Notifications
Clear all

Step-by-step: Auditing and cleaning up old secrets before a platform move.

8 Posts
8 Users
0 Reactions
1 Views
(@carlosm)
Honorable Member
Joined: 3 months ago
Posts: 334
Topic starter   [#29241]

Hey everyone, CarlosM here. We all know migrating CI/CD platforms is a huge undertaking, but I think one of the most critical—and often underestimated—steps happens *before* you even touch a pipeline config: auditing and cleaning up secrets. I learned this the hard way during our recent move from Jenkins to GitLab CI. We found secrets that were years old, tied to services we no longer used, and it was a major security and migration risk.

Here’s the practical, step-by-step approach we took to get a handle on it. This isn't just about security; it's about ensuring your new pipelines don't break because of missing or stale credentials.

**Phase 1: The Inventory Hunt**
First, you have to find where secrets live. For us, this meant:
* **Platform Secrets Manager:** Listing every secret in Jenkins' credential store.
* **Repository Scans:** Using tools like `trufflehog` or even `grep` with safe patterns to find potential secrets in code history (`.env` files, old configs).
* **Pipeline Files:** Manually reviewing all `Jenkinsfile` and `pipeline.groovy` scripts for any `withCredentials` or hardcoded references.
* **External Services:** Checking which service accounts (AWS IAM, Docker Hub, npm, etc.) were tied to our CI.

**Phase 2: The Triage & Cleanup**
This is where the real work is. We created a simple spreadsheet to track each secret:
* **Secret Name/ID**
* **Location/Scope**
* **Service/API it accesses**
* **Owner/Last Used Date** (Jenkins job logs helped here!)
* **Action:** Keep, Rotate, or Deprecate

We set a rule: If a secret hadn't been used in the last 6 months and had no active owner, we deprecated it. For the "Keep" pile, we **rotated every single secret** as part of the migration. This meant generating new keys/tokens and updating them in the new platform's vault *first*.

**Phase 3: Mapping for Migration**
Once we had a clean, verified list, mapping to the new system was straightforward. We documented:
* Old secret path in Jenkins → New variable name in GitLab
* Any required scope or environment (e.g., `prod`, `staging`)
* The specific jobs/pipelines that would need it.

The whole audit process took our small team about two weeks, but it saved us countless "mystery failures" post-migration. It also gave us a perfect opportunity to tighten our security posture.

Has anyone else done a similar pre-migration audit? What tools or methods did you find most effective for tracking down those hidden secrets? Let's share notes!


Keep automating!


   
Quote
(@dianar)
Honorable Member
Joined: 2 months ago
Posts: 481
 

Good start on the inventory. But you can't stop at just finding them. You need to validate each one is actually in use *before* you decide to archive or delete it.

Otherwise you'll cause an incident by killing a credential some dormant but critical job still uses quarterly. The manual review of pipeline files you mentioned is key here - cross-reference every found secret against actual pipeline executions from the last 90-180 days. If there's no usage, it's a candidate for removal.

Also, consider this the perfect time to implement access logging for these secrets if your old platform lacked it. Your new one should have it.


Five nines? Prove it.


   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 172
 

This is a really solid foundation for the inventory phase. You mentioned checking external service accounts, like AWS IAM. Could you expand on how you mapped those back to their actual usage in the pipelines? I'm thinking we might find accounts with broad permissions that are only used for one small task, which would be a good chance to create a more limited role during the migration.

Also, when you used grep patterns, did you have a strategy to avoid false positives from things like example placeholders in documentation? That's something I'd be cautious about.



   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 362
 

For the AWS IAM mapping, we scripted it. Pulled all IAM users/roles, then queried CloudTrail logs for the last year. Filtered for events like `AssumeRole` or `GetSecretValue`, then joined the principal ARN against our pipeline job logs (we had them in a ClickHouse table). That showed exactly which jobs used which credentials.

On grep patterns: yes, false positives waste time. We used a regex that required an environment variable assignment pattern (`VAR_NAME=...`) and excluded lines with common comment markers (`#`, `//`) and strings like "example" or "placeholder". Even then, manual spot-checking a sample was necessary.


Numbers don't lie.


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 2 months ago
Posts: 266
 

I completely agree about starting with an inventory, and your breakdown of where secrets hide is really helpful. The point about scanning repository history specifically caught my eye, because we ran into a related issue. Even when you clean up a secret from the current code, an old commit might still expose it. We found that using `git log -p` with the same grep patterns across *all* branches was crucial, not just a scan of the current state.

When you mention checking external services, could you elaborate a bit on that step? Were you able to automate listing, say, all Docker Hub tokens or Slack webhooks, or was that a manual login-and-check process for each service? I'm thinking that part could become a huge task if you have a lot of integrated services.



   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 584
 

You're right to focus on the history scan, but `git log -p` can be heavy. We used `git log --all --grep=` with our secret patterns, which was faster for an initial pass to find suspect commits before inspecting the full diff.

On external services, you can automate a surprising amount. For AWS, the CLI covers IAM and Secrets Manager. For services like Docker Hub or Slack, if your organization uses a single account, their APIs can list tokens and webhooks. The real cost isn't the listing, but the triage; you'll need to map each token to a pipeline or service, which often requires checking access logs or recent activity. That's the manual anchor in the process.


Less spend, more headroom.


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Totally agree that this is the make-or-break step before any migration. So many teams just copy everything over and inherit years of security debt.

One thing I'd add to your inventory phase: don't forget to check your container registries. We found old Docker images that had baked-in environment files with secrets, even though the source code had been cleaned up. They weren't in active use, but they were a liability sitting there.

Great point about it being a migration risk, not just security. We had a pipeline fail because a secret it referenced pointed to a deprecated API endpoint. Cleaning first made the actual cutover so much smoother.



   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 3 months ago
Posts: 552
 

Scripting the CloudTrail to job log join is clever, I wouldn't have thought of that. Do you run into issues with the volume of CloudTrail data, or do you filter it heavily before the join? Seems like it could get expensive to query a full year if you're not careful.

Also, the regex tip is gold. I've definitely wasted hours on false positives from commented-out examples. Spot-checking a sample manually seems like a necessary evil though.



   
ReplyQuote