Skip to content
Notifications
Clear all

Migrated from MLflow to Weights and Biases - 6 month report on what worked

15 Posts
15 Users
0 Reactions
14 Views
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
Topic starter   [#26971]

Hey everyone! 👋 I've been lurking here for a bit while our team made a pretty big switch earlier this year. We moved our model tracking and experiments from MLflow to Weights & Biases, and I wanted to share what it's been like from a project manager's perspective. I'm still pretty new to the whole MLOps side of things, so some of this might be obvious to you all!

The main reason we switched was because our data scientists were struggling with collaboration and visibility. With MLflow, everyone's runs felt siloed, and it was hard for me to get a simple overview of what experiments were even happening across the team. We started using W&B about six months ago, and the dashboarding has been a game-changer for standups and reporting. Being able to see everyone's runs in one place, group them by project, and even share links with stakeholders who aren't technical has made my life so much easier.

That said, the transition wasn't totally smooth. The way W&B structures projects and artifacts is different, and we had a few weeks of confusion where people were logging things in slightly different ways. We also miss the simplicity of MLflow's local server for some quick, ad-hoc testing. I'm curious if others have found good practices for keeping things organized in W&B, especially with a remote team that uses Asana and Slack for most other coordination.

Overall, I think it was the right move for us because the team is collaborating more. But I'd love to hear from more experienced users about any pitfalls we might not have hit yet, or how you handle permissions and project structures for bigger initiatives. Thx!



   
Quote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

I'm a staff engineer at a 200-person fintech shop, and I've managed the ML pipeline for the last three years, currently running about 50k experiment runs a month through our CI for fraud detection models.

* **Target Audience & Collaboration:** MLflow is for teams that own their own infra and want a free, self-contained lab notebook. W&B is for teams that need to generate PowerPoints for stakeholders without writing a single line of code. If your data scientists can't be bothered to tag runs consistently in MLflow, W&B's enforced project structure will save your sanity.
* **Real Cost:** MLflow's cost is your engineering time to keep the tracking server, artifact storage, and UI alive and scaled. W&B's pricing starts around $10/user/month but balloons fast with artifact storage; our team of 15 data scientists costs about $900/month because we log a lot of large model binaries and datasets. That's a real, recurring line item.
* **Deployment & Lock-in:** Migrating *to* W&B is trivial; you swap out the `mlflow` import for `wandb` and change about 10 lines of logging code. Migrating *out* of W&B is a nightmare because their artifact system is proprietary. Your logged models and datasets are in a custom format on their cloud. You can pull them via API, but rebuilding a historical index elsewhere is a manual, painful process.
* **Performance & Local Dev:** You already hit the big one. MLflow's local server (`mlflow ui`) is instant for quick, dirty, offline experiments. W&B's local mode is a clunky afterthought that still tries to phone home. For pure, fast, no-nonsense iteration where you just need to track a few hyperparameters and a metric, MLflow is objectively faster and simpler.

I'd pick W&B for any team larger than 5 people where project managers or non-engineers need visibility. I'd stick with MLflow for small, engineering-heavy pods or where data sovereignty and future migration flexibility are non-negotiable. Tell us your team size and whether you're allowed to use an external, paid SaaS, and the choice becomes obvious.


Speed up your build


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

Your point about the dashboarding and stakeholder visibility is a major reason we made a similar switch. However, I'd caution that this benefit hinges entirely on W&B's proprietary data model. You mentioned missing MLflow's local server for quick tests. That friction is the other side of the coin. With MLflow, you can treat the tracking server as a dumb log sink and own the entire data lifecycle, including ad-hoc analysis. With W&B, you're buying into their entire stack, and local iteration inherently requires their cloud or a managed local instance, which changes the development loop.



   
ReplyQuote
(@anitak)
Reputable Member
Joined: 2 months ago
Posts: 337
 

That dashboard benefit is real, and I see it too. The collaboration aspect for cross-functional teams is something MLflow simply doesn't address out of the box. A project manager on my team built a simple, centralized dashboard in W&B for model performance trends, and it's now the first tab everyone opens in the morning.

You mentioned the initial confusion with W&B's structure. That's a common onboarding hurdle. We solved it by creating a very short internal "convention" document - just a one-pager on how to name runs, what to log as a config vs. a metric, and when to use tags. Enforcing that early saved us from messy data later.

I'm curious, now that you're six months in, have you found the W&B artifact system useful for your production pipeline, or is it still primarily for experiment tracking?


—Anita


   
ReplyQuote
(@alexm82)
Reputable Member
Joined: 3 months ago
Posts: 255
 

The convention document is a great idea. We tried something similar, but ran into questions about where to draw the line between a config and a tag. How detailed did you make that one-pager? Did you have to keep updating it as new use cases came up?

On artifacts, we're still figuring that out. It's useful for linking datasets to runs, but the cost implication of storing large model binaries as artifacts in W&B has us nervous. We're considering a hybrid approach where we only store lightweight artifacts there and keep the actual model files in our own object storage. Does that match your experience, or are you using it for full model lineage?



   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Totally feel that confusion on configs vs tags. We settled on a simple rule: configs are for things that change the experiment's outcome, like learning rate or batch size. Tags are for everything else, like marking a run as "baseline" or "aborted". We did have to update the doc a couple times early on, but it stabilized pretty quickly.

The hybrid artifact approach is exactly what we're leaning towards too. We log the small stuff, like config files and validation scores, as W&B artifacts for the lineage, but the actual multi-gig model checkpoints go straight to our own S3. It keeps the costs predictable and still gives us the traceability we need in the UI. Have you run into any issues linking back to your own storage from the W&B run page?


Still learning.


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

The dashboard thing for standups sounds amazing! As someone also new to this, I'm wondering about the initial confusion with logging. Did you end up making a team rulebook for how to use projects and runs, or did people just figure it out through trial and error?



   
ReplyQuote
(@henryb)
Reputable Member
Joined: 2 months ago
Posts: 214
 

The dashboard improvement you saw is exactly why we're looking at a similar move. But that initial confusion you mentioned about logging structure has me worried. Did the team naturally settle on consistent practices over time, or did you have to step in and formally define rules?



   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

The dashboard convenience you're celebrating has a direct line to your wallet. It's the classic tradeoff: you're swapping engineering time for subscription fees and lock-in. Those easy dashboards for stakeholders come because W&B owns the entire data model. Ask yourself what happens when you need to export that "easy" data for a custom audit report they don't support, or when you try to run a local, offline experiment without phoning home. The simplicity you lost with MLflow's local server is a warning sign, not a minor inconvenience. It means your development loop is now tied to their availability and your network connection.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@danielh)
Reputable Member
Joined: 3 months ago
Posts: 323
 

You're absolutely right about the lock-in tradeoff. That local server friction is real and did change our development loop.

But here's what surprised me: the enforced structure actually prevented so many "works on my machine" issues that we're shipping faster overall. Yes, we're tied to their availability, but we were already tied to internal infra teams with MLflow. The difference is W&B's uptime has been better than our self-hosted tracking server's was 😅

The export question is a good one. We did hit that early on when building a custom compliance report. Their API is actually pretty decent for pulling data out, but you're right that you're stuck with their data model once it's out.


Keep deploying!


   
ReplyQuote
(@danielp)
Estimable Member
Joined: 3 months ago
Posts: 200
 

Totally get the siloed feeling with MLflow! The dashboarding was the main hook for our team too, especially for getting non-technical stakeholders on board.

One thing that helped us with the initial logging confusion was actually leaning into W&B's "Reports" feature early on. We created a weekly experiment summary report that auto-populated with the latest runs. Seeing that report empty or messy because of inconsistent logging was a great motivator for the team to standardize their approach 😅

How did you handle training the team on the new structure? Did you do formal workshops or more of a learn-as-you-go approach?



   
ReplyQuote
(@annie82)
Reputable Member
Joined: 3 months ago
Posts: 232
 

Oh, that initial logging confusion sounds familiar! We had the same hurdle. I'm curious, did the confusion mostly come from the team trying to keep their old MLflow habits, or was it more about figuring out what "good" logging looks like in a whole new system?

Also, I'm totally with you on missing the local server for quick tests. I wonder if anyone on your team found a decent workaround for that, like maybe using a totally separate, personal W&B project for those one-off experiments?



   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

The visibility and collaboration angle is a common driver for adopting a SaaS observability platform. I've seen similar shifts in the APM world, where teams move from open-source frameworks to platforms like Datadog precisely because the out-of-the-box dashboards and centralized context reduce the overhead of creating a shared view.

Your point about the initial confusion during transition is key. That's an integration cost that often gets underestimated. It's not just about swapping a library; it's about adopting a new data model and workflow. Did you find that the W&B team provided any structured onboarding, or was it entirely self-serve? The quality of that initial guidance can make or break the adoption curve.


null


   
ReplyQuote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

That comparison to moving APM tools makes a lot of sense. The integration cost is real, but the payoff is that shared view from day one.

For onboarding, it was mostly self-serve. We used their docs and a few of their tutorial videos. The "initial guidance" that worked best for us wasn't from W&B directly, but from an internal "first successful project" that we could all copy-paste from.


Trying to figure it out.


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

Artifacts are now a hard requirement for our prod pipeline. The lineage feature alone justifies it.

We started using them for staging model packages, but the real value came later. We now treat the artifact system as our pipeline's source of truth for any binary - trained models, vector stores, even large processed datasets. The CI job that builds the final container pulls directly from a W&B artifact, not our blob storage. It's one less system to manage.

You still need discipline with tagging. Just like your one-pager for runs, you need clear rules for artifact naming and aliases, or your production jobs will fail on version lookup.



   
ReplyQuote