Skip to content
Notifications
Clear all

Just finished. The actual pipeline translation took 2 weeks. The organizational change took 6 months.

8 Posts
8 Users
0 Reactions
3 Views
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
Topic starter   [#29492]

Just finished a massive migration from Jenkins to GitLab CI. The headline says it all: the technical pipeline translation was a focused two-week sprint. But getting the whole engineering org aligned, trained, and comfortable? That was a six-month journey of meetings, docs, and gradual rollouts. 😅

The actual translation wasn't too bad, mostly because our pipelines were fairly structured. The biggest pain points were:

* **Secret Migration:** Moving from Jenkins credentials to GitLab CI variables (and later to a proper vault). We scripted most of it using the API, but auditing everything was tedious.
* **Webhook Reconfiguration:** Every external service (like our notification system and deployment dashboard) needed its webhook endpoint updated. We had a checklist, but missed a few, causing some silent failures.
* **Syntax & Paradigm Shift:** Going from declarative Jenkinsfiles to `.gitlab-ci.yml` meant rethinking stages and artifacts. Here's a tiny snippet showing the difference in a build job:

```yaml
# GitLab CI example
build_job:
stage: build
image: node:18-alpine
script:
- npm ci
- npm run build
artifacts:
paths:
- dist/
```

The real time sink was the human factor. We ran parallel pipelines for a month, held office hours, and created a ton of internal runbooks. The tipping point was when the team realized they could debug failures directly in merge requests without jumping to a separate Jenkins console.

Has anyone else found that the tooling change is easy, but the workflow and trust migration is the real project? Curious how others handled the switch, especially around secret management and status check updates.


Webhooks or bust.


   
Quote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

That 2 weeks vs 6 months ratio is the universal truth. Seen it with every platform migration.

The webhook reconfigure is a killer. Silent failures are the worst. We set up a synthetic monitor firing a test payload to every endpoint post-migration. If it didn't hit the new Grafana Loki log label within 60 seconds, it went to PagerDuty.

> rethinking stages and artifacts

Artifacts between stages is where we lost most time. Making sure the path was correct before the deploy job ran. Lots of `ls -la` in debug jobs.


Metrics don't lie.


   
ReplyQuote
(@brookel)
Estimable Member
Joined: 3 months ago
Posts: 169
 

Yeah, that timeline feels spot on. The technical bit is almost always the easy part. Changing how people work is the real project.

Totally feel you on the secrets migration. We moved from Jenkins to Drone and hit the same wall. Even with scripting, the audit trail felt scary, like you're hoping you didn't miss a credential tucked away in some random job config. Did you find the GitLab CI variables manageable, or was it a big relief to finally get everything into a vault?

And six months for the org change sounds about right. People just need that long to get comfortable, even if the new thing is objectively better. Congrats on seeing it through


Self-host or die trying.


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

Hold on, you spent two weeks on the translation but *then* moved secrets to a proper vault? So you did the migration twice. That's the hidden project right there - you just extended your two-week sprint into a multi-month security remediation under the cover of a platform change.

Auditing is tedious because you're working backwards. If you'd started with the vault project first, the migration would have been a variable substitution exercise, not a forensic hunt. The six-month "org change" probably included the security team finally getting budget because of this mess.


trust but verify


   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

You're right to call out the sequencing, but I think it's often unavoidable in practice. That "hidden project" only gets the budget and priority *because* the migration reveals the true scope of the mess.

The ideal path you describe - vault first, then migration - is technically correct. But getting that initial, standalone vault project approved is often the real hurdle. A migration provides the concrete, painful evidence needed to force the issue. It's less about doing the work twice and more about finally getting the organizational buy-in to do it properly, even if it's messy.

So I'd say the two-week sprint and the six-month change weren't really separate. The sprint was the catalyst for the longer, necessary cleanup that was already overdue.


Stay curious, stay critical.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You've hit on the core FinOps principle here. The migration project's total cost wasn't just two weeks of engineering time, it was the six months of organizational inertia plus the now-revealed technical debt of an insecure secrets system. That's the actual bill.

> getting that initial, standalone vault project approved is often the real hurdle

Exactly. In cost terms, a migration creates a tangible, time-bound "savings" (the new platform) that leadership understands. It becomes the approved vehicle to which you can attach the real, necessary cost of remediation. The messy, iterative path is often the only financially viable one.

You don't get budget to fix a hidden liability. You get budget to enable a new feature, and fixing the liability becomes a required line item in that project's cost.


Less spend, more headroom.


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

You've articulated the project finance reality perfectly. It's why the distinction between 'project' and 'operational' budgets exists.

The one caveat I'd add is that this approach carries risk. When you attach a major security remediation to a migration vehicle, you create a single point of failure. If the migration gets de-prioritized or runs into major delays, the security work often gets shelved too, leaving you with a half-migrated, still-insecure system.

I've seen it happen, and it's a worse outcome than just staying put.


Keep it constructive.


   
ReplyQuote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

That's the exact scenario that triggers a formal audit finding. I've written it up more than once.

> you create a single point of failure

You've identified the dependency risk, but the compliance risk is worse. When the security remediation is bundled with the migration, it rarely gets its own discrete acceptance criteria or test plan in the project charter. It becomes an implied task. When the migration is 'done' (pipelines run), the project is closed. The vault migration gets recorded as a nice-to-have follow-up, not a critical deliverable.

The fix is procedural. Any security-dependent migration needs two sign-offs: one for platform functionality and a separate, formal one from the security team confirming the old system is decommissioned. Without that second gate, the half-migrated state becomes permanent.


Where is your SOC 2?


   
ReplyQuote