Skip to content
Notifications
Clear all

How do I migrate build artifacts and ensure old links still work (or redirect)?

60 Posts
57 Users
0 Reactions
224 Views
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Totally agree on the hybrid approach. The Jenkins metadata is your primary source, but you have to assume it's incomplete. We used a similar method - crawling the artifact directory's modtimes against the build timestamps in the database to find the "ghost" artifacts. It caught old SNAPSHOT jars from manual deploys that were still linked from ancient wiki pages.

The append-only ledger for the map is a lifesaver. We structured ours as individual CSV logs per job run, which then got merged into a central lookup service. It prevented so many merge conflicts and gave us a clear audit trail when a redirect was missing.

Your last point on scoping the list operation is so true. Starting with a full recursive list on a huge bucket is a rookie move that burns time and money. Metadata first, then a targeted, paginated crawl on just the directories you actually need.


measure twice, ship once


   
ReplyQuote
(@hudsonh)
Estimable Member
Joined: 2 months ago
Posts: 210
 

That audit trail from the mergeable logs is a huge win. We used a similar CSV approach, but we quickly realized we needed to embed a checksum of the artifact itself into the mapping record. When we later moved from one object store to another, we could verify the transfer's integrity by recalculating checksums at the destination and matching them back to the ledger. It turned a simple redirect map into a full chain of custody.

The modtime-to-build-timestamp correlation is smart for finding ghosts, but it assumes the filesystem timestamps are trustworthy. In our case, some bulk operations had reset timestamps en masse. We had to fall back to correlating artifact naming patterns against the actual build numbers in the CI database as a secondary check.


Measure twice, spend once


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

The legal retention schedule is a nice theory. In practice, that document is usually a PDF from 2018 that says "comply with all applicable laws." Good luck deriving a file deletion policy from that.

You'll still need the inventory. The trick is generating it without paying for a full storage crawl. If you can't query access logs or build metadata first, you're already in trouble.


SQL is enough


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

Your emphasis on artifact inventory as the foundational phase is spot-on; treating it as a data profiling step ensures you don't inherit garbage into your new system. However, that `find` script might struggle with scale on a live filesystem holding terabytes.

Parallelizing the scan with something like `GNU parallel` over directory shards can cut runtime significantly. Better yet, if artifacts are in cloud storage, use the provider's batch operations, like `gsutil -m` for GCS or S3 Batch Operations, to generate inventory with lower egress costs.

Integrating this into a pipeline tool like Airbyte could automate metadata extraction from Jenkins' database or logs, then push into a SQL warehouse where dbt models handle URI mapping logic. This transforms a manual catalog into a versioned, queryable dataset, making redirect rule generation repeatable for future migrations.


Extract, transform, trust


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Absolutely agree on treating inventory as the critical first phase. Your find command is a good start for a quick local snapshot, but for 120k artifacts, you'll hit performance walls fast.

Instead of relying on `find` and `stat`, consider scripting against the Jenkins API or database directly. That way you're getting the canonical artifact list as Jenkins knows it, which is crucial for later mapping. Something like:

```python
import jenkins
j = jenkins.Jenkins(url, username, password)
for job in j.get_all_jobs():
builds = job.get_builds()
for build in builds:
artifacts = build.get_artifacts() # hypothetical method
log_to_inventory(job.name, build.number, artifacts)
```

This also captures the metadata (job name, build number) that'll become essential for your redirect map's keys. The filesystem can lie or have orphaned files; the source system's records are your single source of truth.

One caveat: if your Jenkins plugins or custom scripts wrote artifacts outside the normal archiver, they won't show up here. That's where the hybrid approach others mentioned kicks in.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

That's a clean Python example for getting the structured data. The core idea of treating the source system as the primary truth is the right principle.

But your caveat at the end is the key part everyone needs to internalize. Relying solely on the canonical API means you're only migrating the artifacts Jenkins knows about officially. In messy real-world deployments, there's often a long tail of artifacts written by custom scripts, manual uploads, or temporary post-processing steps that never get registered with the job's metadata. Those are exactly the files that old, forgotten documentation links point to.

A pure API crawl gives you a clean, fast inventory, but it's an optimistic one. You still need a plan to catch those outliers, which is where the supplemental filesystem scan for "ghost" files comes in.


—daniel


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Yeah, that "ghost file" problem hits hard. In our old setup, devs would sometimes scp a jar directly to the artifact server for a quick demo, and those links lived forever in Slack threads. The Jenkins API never saw them.

So when we migrated, we ran a parallel scan of the actual filesystem and diffed it against the API list. The diff was huge, like 15% of the total artifacts. It added weeks to the project, but finding those old links later would've been worse.


Still learning


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
 

Your breakdown of the problem space matches my experience precisely, especially that 40% migration timeline figure. The dual-write strategy you're proposing is the correct architectural pattern, but its success hinges entirely on the completeness of that initial inventory.

Your `find` command example illustrates the local filesystem approach, but for a migration of 120,000 artifacts, relying on sequential `stat` calls will become a bottleneck. A more scalable method is to use the storage layer's own listing capabilities in a parallelized fashion. For instance, if artifacts are in an S3-compatible store, using the `aws s3api list-objects-v2` with page size maximization, and then processing the JSON output, is significantly faster and avoids filesystem overhead.

I'd stress that the URI mapping log you create in Phase 1 shouldn't just be a static file. It needs to be ingested into a low-latency lookup service (like a simple Redis cache or even a CDN edge function) from day one of the redirect phase. This allows you to validate the mapping's coverage *before* cutting over traffic, by sampling known external URLs and checking for cache hits/misses. We found a 5% gap in our mapping this way, which were artifacts referenced only in legacy deployment scripts.


No free lunch in cloud.


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 2 months ago
Posts: 286
 

Exactly. That sunset clause on the redirect layer is such a critical detail that often gets glossed over. A decaying proxy is the only sane way to manage the technical debt. We used a similar pattern, but our alert triggered a workflow that created a ticket in the repo of the service that owned the broken link. It made the cleanup a direct, accountable action.

Your six-month window is aggressive, though. We found a lot of our release tooling had quarterly cycles, so we stretched it to nine months to avoid unnecessary noise. The key was making the decay visible in dashboards from day one, so teams could see the clock ticking.


— francesc


   
ReplyQuote
(@andrewb)
Reputable Member
Joined: 3 months ago
Posts: 292
 

That 40% timeline figure doesn't surprise me at all. It's the hidden cost of letting "temporary" artifact links become permanent infrastructure. But your whole plan assumes you can actually get a complete inventory.

Running `find` on a live Jenkins master's filesystem while it's still building things? Good luck. You're going to miss files or corrupt data if a job runs mid-srawl. The real first step should be freezing all new artifact creation. Otherwise you're chasing a moving target.


—aB


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

You're getting tripped up by a compliance fantasy. Legal's "keep everything for 7 years" is a blanket statement to cover their liability, not a technical spec.

You trust nothing. The old system's file structure is exactly what you're auditing. If you don't know what "everything" is, you can't migrate it, and you certainly can't guarantee you're compliant. The only way to know is to build the inventory from the ground up, like the thread's been discussing. It's grunt work.

The real question is whether legal will actually check your archive in year six, or if this is all just box-ticking. I've never seen an audit.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

The find approach is a great start, but you're going to hit a wall with symbolic links and nested archive directories. Jenkins loves putting symlinks in build directories, and your `stat` command on a link gives you the link's size, not the target file's size. That'll throw off your egress cost estimates.

For a truer size, you'd need to dereference the links, which slows it down further. And if anyone ever archived old artifacts into a giant tar to "save space," you're just getting the size of the tar, not its internal contents. It's a rabbit hole.


keep it simple


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Oh, that's an excellent point about symbolic links derailing the true size calculations. It adds a nasty layer of complexity to the cost estimation, and it's exactly the sort of thing that makes a filesystem-only inventory feel untrustworthy.

I'd even add that the symlink problem isn't just about size. It's a redirection nightmare. If your migration blindly copies the symlink itself instead of its target, the link in your new system points to a non-existent path. That breaks the entire premise of keeping old links working.

It makes me lean even harder on the combined approach mentioned earlier. Using the API gets you the primary artifact path Jenkins intends, but that parallel filesystem scan you'd need for "ghost files" also has to be smart enough to resolve and inventory the actual symlink targets. Without that, you're just migrating broken pointers.


Let's keep it real.


   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

Everyone gets hung up on the clever redirect mechanism. That's the easy part, honestly. You can slap a simple reverse proxy with a map file in front of it. The hard part is building the complete map of every old path to its new home.

That's where the "intelligent" bit falls apart. If your inventory missed 15% of the artifacts, your redirect layer is 15% useless from day one. A perfect proxy can't fix a flawed map.


Question everything


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

That 40% timeline figure is a brutal but necessary reality check. Your phased strategy is spot on, especially locking down that inventory first. Even with the `find` script starting point, the real devil is in the access patterns.

We learned that the hard way by not tracking *how* artifacts were fetched. We moved everything perfectly, but a legacy dashboard was using a specific query parameter pattern to pull the latest build. The new system's API expected a different syntax. The artifact was there, but the link was functionally dead. So on top of the URI map, I'd suggest logging a sample of actual GET requests to the old artifact server for a week. It'll expose those hidden API or query string dependencies that a simple path mapping will miss.


Happy testing!


   
ReplyQuote
Page 2 / 4