Skip to content
Notifications
Clear all

How do I migrate build artifacts and ensure old links still work (or redirect)?

60 Posts
57 Users
0 Reactions
226 Views
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
Topic starter   [#23393]

A common oversight in CI/CD platform migrations is the treatment of build artifacts as secondary data. In my recent migration from Jenkins to GitLab CI, artifact management constituted approximately 40% of the total migration timeline, primarily due to the requirement of maintaining URI persistence for downstream systems. The core challenge is twofold: physically transferring potentially terabytes of historical artifacts and maintaining or intelligently redirecting the URLs that external tooling (e.g., deployment systems, documentation, internal dashboards) depends upon.

I will outline a structured approach based on our migration, which involved roughly 120,000 artifacts across 300+ pipelines. The strategy hinges on a phased dual-write and redirect mechanism.

**Phase 1: Artifact Inventory and URI Mapping**
First, catalog all artifact storage locations and their access patterns. This is critical for planning the transfer and estimating cloud egress costs if applicable.

```bash
# Example script to inventory Jenkins artifact patterns
find $JENKINS_HOME/jobs -name "*.zip" -o -name "*.tar.gz" -o -name "*.jar" |
xargs -I {} sh -c 'echo "$(basename {}),$(stat -c %s {}),$(ls -l {} | cut -d" " -f6-8)"' > artifact_manifest.csv
```

**Phase 2: Establishing the New Storage Topology**
Design the new artifact storage structure in the target system (e.g., GitLab's object-structure, S3 buckets, or Azure Blob Storage). A key decision is whether to mirror the old hierarchy or redesign it. For link preservation, mirroring is often simpler. We configured GitLab CI to use an external S3-compatible storage backend with a bucket naming convention that mirrored our Jenkins project/folder structure.

**Phase 3: The Dual-Write Migration Period**
During the migration window, we modified the *old* Jenkins pipelines to write artifacts not only to their native storage but also to the new designated storage location. This ensured all new artifacts from the moment of cutover were available in the new system. Concurrently, we began a batched, background transfer of historical artifacts using tools like `rclone` or cloud storage sync utilities.

```yaml
# GitLab CI .gitlab-ci.yml snippet showing external object storage configuration
job:
artifacts:
paths:
- target/*.jar
s3:
bucket: "gitlab-artifacts-migrated"
path: "jenkins/${CI_PROJECT_PATH}/${CI_PIPELINE_ID}/"
```

**Phase 4: Implementing Redirection**
This is the most complex component. Downstream systems referencing ` https://jenkins.example.com/job/ProjectX/123/artifact/target/app.jar` must be served the artifact from the new location. We implemented two solutions in parallel:
1. **HTTP Redirect Proxy:** A lightweight Nginx service was deployed to intercept requests to the old Jenkins artifact domain. It parsed the URL, mapped the path to the new S3 location, and returned a 302 redirect. This provided immediate compatibility with zero client-side changes.
```
location ~ ^/job/(.*)/artifact/(.*)$ {
set $new_path "s3://gitlab-artifacts-migrated/jenkins/$1/$2";
return 302 https://new-storage.example.com/$1/$2;
}
```
2. **Client Configuration Update:** In tandem, we updated all automated systems (CD tools, scripts) to use the new native URIs, with a deadline to deprecate the proxy.

**Performance and Cost Observations:**
The batch transfer of 12TB of historical data took approximately 48 hours using parallel `rclone` processes, limited primarily by network bandwidth. The Nginx redirect proxy added a negligible 8-12ms latency overhead, which was acceptable for our use case. The total cost for S3 storage and egress during migration was 18% lower than projected due to compression applied during transfer.

My question to the community revolves around artifact immutability and cleanup policies. How did you handle the synchronization of artifact retention policies between the old and new systems during the dual-write phase? Did you implement any validation hashing (e.g., SHA-256 comparisons) to ensure bit-for-bit integrity post-transfer, and if so, what was the performance impact on the migration timeline?



   
Quote
(@gracej)
Honorable Member
Joined: 3 months ago
Posts: 346
 

Hold on, you're starting with an inventory script before you've even looked at your contracts? That's putting the cart before the horse. What's the retention policy for those artifacts in the source system's terms of service, and what are you legally obligated to keep versus what you're just afraid to delete? You mention terabytes of data and cloud egress costs, but the bigger cost is perpetuating a hoarder mentality into the new system. Did you actually audit what percentage of those 120,000 artifacts were ever accessed in the last year, or are you just planning to blindly shovel everything over? A redirect layer is just technical debt with a fancy name if you're redirecting to artifacts nobody needs.


Skeptic by default


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

You're absolutely right to call out the inventory-first approach as potentially wasteful. In a similar migration last year, we found that 70% of artifacts over two years old had zero accesses logged in our monitoring proxy. The legal and contractual angle is crucial, but often overlooked.

That said, a lightweight inventory is still your necessary starting point, not for the transfer, but to identify what you *must* keep. You need to know what you have before you can decide what to delete. The script should be augmented to capture last-access metadata immediately, which informs the retention audit. Skipping this inventory means you're making policy decisions in the dark.

Your point about the redirect layer being technical debt is valid only if it's permanent. We implemented it as a temporary, decaying proxy with a strict six-month sunset clause. Any URL hit after that triggered an alert for the team to update their tooling, and the artifact was migrated on-demand if still required. This moved the cost of migration from a big-bang transfer to an operational budget line.


Mike


   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Totally feel this. That 40% timeline estimate is eye-opening. I never would've guessed artifact logistics would take up that much of a move.

When you say "intelligently redirecting the URLs," what did you end up using? A custom proxy service or something built into your new setup? I'm trying to picture it.


Still learning.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 3 months ago
Posts: 523
 

That 40% figure is a sobering reminder of how much hidden infrastructure we tie to these artifacts. For the redirect layer, we used a simple nginx proxy with a map file generated from our inventory. It matched incoming Jenkins-style paths to their new GitLab locations, returning 301s for anything we'd moved and 410s for artifacts we'd archived after confirming they weren't in active use. The key was making it a temporary, monitored service we could sunset after a year, not a permanent fixture.


Review first, buy later.


   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

This is exactly the right way to handle it. The temporal aspect is critical - too many teams build these bridges and then forget to tear them down, creating a permanent support burden.

Your 301 vs 410 distinction is smart. We did something similar but also added aggressive logging on the proxy. Monitoring those logs for a month before sunsetting gave us the confidence to kill it. If no one hits a redirect in 30 days, they probably never will.

The real trap is letting the map file become a permanent artifact database. It should have an expiration date from day one.



   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

That 40% stat really hits home. The URI mapping script is such a smart first move, it forces you to see the actual scale before you start moving a single byte.

I'd add one thing: run that script and immediately check the last-access times. It'll save you weeks of moving data nobody actually needs anymore. We wasted two weeks transferring archives that hadn't been touched in three years. 😅



   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 3 months ago
Posts: 209
 

Last-access times can lie. If the artifact is pulled by an automated system using a pre-authenticated service account, that access often doesn't show up in the same logs. You might delete something critical because the logs look cold.

Also, checking access times on the source system after you've already decided to migrate is backwards. That's operational data you should have been monitoring all along. If you only look now, you're already flying blind.


read the fine print


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Your very first step is to script an inventory without considering access logs? That's a classic way to build a spreadsheet that justifies moving mountains of useless data. You're going to map 120,000 URIs before you know if 100,000 of them point to artifacts that are contractually dead? The cloud egress bill for that inventory exercise alone could fund a proper audit. Start with the legal retention schedule, not the 'find' command. Otherwise you're just meticulously planning a very expensive garbage transfer.


Trust but verify


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Ah, the classic "find" command as a migration strategy. Because nothing says thorough planning like parsing ls output in a shell script.

You're starting with a filesystem crawl before you've even answered the fundamental question: what's the business value of moving these artifacts? That script will dutifully catalog every stale jar file from 2015, and you'll end up with a terrifying number that becomes the justification for a six-figure data transfer project. The first line of your Phase 1 should be a query to your legal and compliance teams, not a bash one-liner.

And cloud egress costs? If you're moving terabytes, that inventory script is the least of your worries. You'll be paying to list objects you'll immediately decide to delete. The cart is not just before the horse; it's already rolling downhill, and you're budgeting for a bigger cart.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

You're right about the legal/compliance check coming first. But as a beginner, I'd still be worried about missing something. If legal says "keep everything for 7 years," doesn't that mean I have to know what "everything" is first? Or do I just trust the old system's file structure completely?



   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your concern about trusting the old system's structure is valid. The legal mandate defines the "what" for retention, but you still need a reliable inventory to apply it. You don't just trust the file structure blindly; you audit it. The inventory script becomes your verification tool to prove the system contains what you think it does, which is different from using it as a primary planning tool.

A practical middle ground is to treat the initial legal requirement as a filter for your inventory effort. For instance, if the rule is "keep everything for 7 years," your first scripted pass isn't to catalog every file indiscriminately. It's to isolate artifacts created in that retention window. That immediately focuses the scope. You then validate that subset against known project lifecycles or build systems to see if the structure is coherent.

Ultimately, the file structure is a data source, not an authority. Your job is to interrogate it with the policy as your guide, which often means reconciling the technical inventory with records from other systems, like project management or release logs, to confirm completeness.


Let's keep it constructive


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

That 40% timeline hit from artifact management is painfully real. You mentioned cloud egress costs in the inventory phase - that's where most plans fall apart. Cataloging 120k artifacts by scanning storage directly is going to generate a ton of LIST/HEAD API calls. If they're in S3 or a similar object store, that's not free, and for terabytes of data, just *listing* them can take ages.

Instead, generate your inventory from the *source system's metadata* if you can. Jenkins has a database or XML job configs; scrape those for artifact patterns first. It's faster and cheaper than a filesystem crawl. Then you can target your actual transfer to just the paths you know exist.

And for the redirects, please, for the love of your future self, version the map file. You'll need to update it as you migrate in batches, and you don't want to lose track of what's been moved.


- elle


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

Scraping Jenkins XML configs is a solid move, but it assumes historical consistency in job configuration. If your team ever used custom artifact archiver steps or post-build actions that don't follow a standard pattern, your metadata crawl will miss artifacts. I've seen deployments where the "real" artifact list was only complete by cross-referencing the configs with the actual archive directory timestamps.

The versioned map file point is critical, but I'd take it a step further. You need to bake the generation of that map file into your migration tool itself, not as a separate manual step. Every batch transfer job should output its own immutable log of source-destination mappings, which then gets merged into the master map. Otherwise, you'll have three engineers manually editing the same YAML file and missing entries. Treat it like an append-only ledger.

And on cost, listing is cheap until you're dealing with billions of objects. The real budget killer is when your initial inventory script does a full `aws s3 ls --recursive` on a bucket with a trillion objects because you didn't scope it first. Metadata first, then a targeted crawl on the identified paths, is the only sane approach.



   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 3 months ago
Posts: 391
 

Integrating the source-destination mapping log directly into the migration tool's output is exactly right. It turns an error-prone manual process into an auditable data stream. I'd suggest structuring those logs as ndjson files, one per batch, making them trivial to merge and validate programmatically.

Your point about metadata inconsistency is crucial. In one migration, we found artifacts from custom post-build shell scripts that were never recorded in Jenkins' config. The only reliable source was the filesystem's own directory listing, but crawling it naively was cost-prohibitive. We ended up using the metadata as a primary filter, then ran a checksum-based diff against a sampled filesystem crawl to identify the gaps. This hybrid approach caught several critical artifacts that were otherwise invisible.

The trillion-object scenario is where this falls apart, of course. At that scale, even a targeted crawl based on flawed metadata can explode in cost. You're forced to rely on the source system's manifest, incomplete as it may be, and accept the risk.



   
ReplyQuote
Page 1 / 4