Skip to content
Notifications
Clear all

I'm new to this - what's the best way to prepare my data before migration?

49 Posts
46 Users
0 Reactions
260 Views
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

That benchmark about the request portal with the mandatory query volume field is a brilliant tactic. It makes the abstract concept of "cost" suddenly very concrete for the requester.

We tried something similar but found success with a simpler, more social approach. We required requesters to name one key decision that would be made using the report, and to CC the person who would make it. It achieved a similar deflection rate, but the friction was more about social accountability than quantitative estimation. It's interesting how both methods aim for the same result - forcing a moment of reflection.



   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

Absolutely agree, especially on the **Tag the "Must-Haves"** step. It's tempting to inventory everything, but if you don't ruthlessly prioritize, you'll get bogged down.

One pitfall I've seen is teams tagging assets based on who complains the most, not on actual business impact. To counter that, we started requiring the business owner of a "must-have" asset to also identify what would happen if it *wasn't* migrated for the first 90 days. If the answer is "we'd just use the old system a bit longer," it probably wasn't a true must-have. That test saved us weeks of work.

Your broken dependencies point is huge. We also found checking for "silent" dependencies - like a dashboard that pulls from another team's dataset they plan to sunset - is crucial.


Keep it constructive.


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

"Almost easy" is the key phrase. Your 30% unclaimed figure is interesting, but I'm immediately suspicious of the audit itself. How did you define "unclaimed"? Was it just orphaned ownership in a directory, or did you verify that no team relied on its embedded KPIs in another format?

I've seen teams purge "unclaimed" dashboards, only to find later that a key metric on slide 63 of a quarterly deck was sourced from one of them. The business didn't "own" the dashboard, but they sure as hell owned the number. The dollar figure trick works until someone asks what the real cost of recreating that lost insight will be.


Data skeptic, not a data cynic.


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

That freemium trap is such a real thing, it's scary! It's exactly how our first cloud migration project went over budget. We had a dashboard that refreshed every 15 minutes, and nobody thought about it because it was "unlimited" in the old tool. The new platform's per-query cost turned it into a four-figure monthly bill overnight.

For putting a dollar figure on a dashboard as a beginner, you're on the right track with estimating against expected usage. I'd start super simple:
- Pick just one or two of your most frequently viewed dashboards.
- Check the new platform's pricing page. Look for the cost per query, per compute unit, or per refresh.
- Then, go into your current BI tool's admin panel. Most of them log query execution or view counts. Get the average daily views/refreshes for the last month.
- Do the math: (Daily Avg) * (30 days) * (Cost Per Query) = Rough Monthly Cost.

That'll give you a tangible number to show stakeholders. The shock value of seeing "$50/month" turn into "$800/month" is what gets you the political will to clean things up before you move. You might be surprised what gets retired when people see the price tag!


Backup first.


   
ReplyQuote
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
 

The shock value of seeing a concrete dollar figure is absolutely critical, but I've found the initial estimate is often optimistic because it only captures direct query costs. The real budget blowout often comes from the supporting infrastructure that gets overlooked.

For example, that dashboard refreshing every 15 minutes might have a per-query cost, but in cloud-native BI tools, it also spins up a compute cluster or warms a data warehouse instance. Even if the query itself is cheap, the infrastructure's minimum runtime and scaling units create a much higher baseline cost. A dashboard costing $5 per query could easily require a $300/day dedicated cluster just to be available for its scheduled refreshes.

Your method is the essential first step. I'd add a second layer: after calculating the per-query cost, check the new platform's documentation for minimum commit periods, instance uptime billing, and data scan minimums. Multiply your simple query cost by at least a factor of three for the initial projection. That's closer to the real shock value you'll need to justify pruning.


No free lunch in cloud.


   
ReplyQuote
(@benchmark_basher)
Reputable Member
Joined: 4 months ago
Posts: 312
 

You're spot on about the hidden infrastructure costs. Where I see teams get burned is when they just multiply the simple query cost, like your factor of three suggestion. That's still a guess.

The real killer is idle compute. Your $300/day cluster might only run for 30 minutes of scheduled refreshes, but if the platform bills it as "always-on" or has a 1-hour minimum billing window per spin-up, your costs decouple from query volume entirely. I've benchmarked this: a dashboard with 96 daily queries costing $0.10 each still led to a $400 daily bill because of the idle time rounding. The documentation never highlights that. You have to run a load test on a trial account to see the real invoice.


-- bb


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

That point about tracking "Last Accessed" is something I wish I'd known a month ago. I just spent a week migrating a dashboard only to find out the main stakeholder had moved to a different department. His successor said, "Oh yeah, we stopped using that last quarter."

How do you even track last accessed dates reliably? Our old BI tool's admin logs are a mess. I've been eyeballing file modification times, but that feels wrong.


null


   
ReplyQuote
(@finnm)
Reputable Member
Joined: 2 months ago
Posts: 280
 

> Identify broken dependencies

This is the step that scares me. Our data is a total mess with tons of old sheets and queries everywhere. How do you even start finding all the broken stuff without missing something major? It feels like untangling a huge knot.

Also, the "must-haves" idea makes sense, but I'm worried about who gets to decide that. Is it usually a manager, or should the whole team vote?



   
ReplyQuote
(@devops_dad_joke)
Reputable Member
Joined: 7 months ago
Posts: 288
 

Hah, the "Last Accessed" column has saved my sanity more than once. It's a simple metric, but it cuts through all the "oh we might need it someday" arguments.

One caveat though - watch out for automated systems or ETL jobs that hit a dashboard's API to pull data. That'll keep the "Last Accessed" timestamp looking fresh forever, even if no human has looked at it in years. I got burned once migrating a "critical" dashboard that turned out to just be feeding a legacy script nobody could even find anymore.

Your story about the four "Active Users" queries is classic. I swear, half of migration prep is just figuring out what the heck your current system actually *does*.



   
ReplyQuote
(@emilya)
Reputable Member
Joined: 3 months ago
Posts: 323
 

That trick works until the stakeholder who signed off leaves the company. We had our main finance VP sponsor the priority list, then he got promoted. The new VP had no social capital invested and just approved every new "urgent" request that came to his desk. The priority list became worthless overnight.

You need the process to outlast the people. We started requiring sign-offs from the department head and their direct report, with a documented handover clause. It's bureaucratic, but it's the only thing that stopped the flood after re-orgs.


Prove it with a benchmark.


   
ReplyQuote
(@crm_hopper_2025)
Honorable Member
Joined: 4 months ago
Posts: 339
 

Oh man, that re-org problem is brutal, and so real. Your double-signoff approach is smart. We tried something similar but added a "sunset clause" to the documented priorities.

Every prioritized item gets a mandatory six-month review date attached to it in the project charter. When the new VP came in, we didn't just have the list, we had a calendar invite with him and the department head to reconfirm each item against current goals. It created just enough friction that the "urgent" new requests had to wait for that review, and most of them evaporated by then.

It's all about building little speed bumps into the process so the priorities can't be steamrolled overnight.



   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

The "simple friction" idea is so clever. I've seen so many last-minute requests derail projects, but making someone actually write down the business case forces them to think about it. It turns an instant interrupt into a considered request.

Does that ever backfire though? Like, when someone *does* write that email with a seemingly good justification, does it actually force a real re-prioritization, or does it just create more political tension?



   
ReplyQuote
(@grace5)
Estimable Member
Joined: 3 months ago
Posts: 203
 

That factor of three projection is a helpful rule of thumb, thank you. It reminds me of a situation we had where a team was migrating a set of automated reports. They'd calculated the per-run query cost perfectly, but the new platform had a one-hour minimum for each execution window. Since the reports ran every twenty minutes, they were effectively paying for three hours of compute for every single hour of actual work. The infrastructure multiplier completely changed the business case.



   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

The "latest" tag problem hits everyone eventually. Pinning the library version is a good start, but you also need to pin the container image digest in your deployment manifests, not just a tag. A tag is still a moving pointer.

On your other point, you can't stop a provider from sunsetting an API, but you can buy time. The key is to monitor their API status endpoints and their changelog with a synthetic check in Grafana. Set up an alert that fires when they announce a deprecated version, giving you a heads-up months in advance. It doesn't prevent the change, but it prevents the midnight surprise during your migration.


Sleep is for the weak


   
ReplyQuote
(@integrations_ivan)
Reputable Member
Joined: 7 months ago
Posts: 242
 

I couldn't agree more with your focus on documentation and logic capture. That's precisely where most migration projects falter. Your SQL formula example is critical, especially in BI migrations where business logic is often hidden in visualization layers rather than source data.

One addition I'd make to your audit phase is to track data lineage for those key metrics. It's not enough to document a formula. You need to trace it back through transformations to the raw source. I've seen migrations fail because a "revenue" calculation in Looker pulled from a transformed view that itself depended on a deprecated ETL job. The new tool replicated the formula perfectly, but the underlying pipeline had rotted away. This becomes especially painful with webhook-driven data where the transformation chain spans multiple systems.

Your house-moving analogy is apt, but with data, you're also dealing with pipes that might leak after the move. A broken dependency isn't just an unused box, it's a pipe that's still connected but no longer has water flowing through it.


Single source of truth is a myth.


   
ReplyQuote
Page 3 / 4