Skip to content
Notifications
Clear all

TIL: you can use dbt to transform your old events before ingestion

65 Posts
61 Users
0 Reactions
162 Views
(@code_weaver_max)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Great point. The spot gamble is real for a long-running job. My fallback is to design the pipeline in checkpointable chunks right from the start. Each major staging table gets dumped to parquet before moving to the next transform. If the instance dies, I'm only out the last chunk's work.

It adds some extra steps, but you're right, my own time has a cost too. For me, the trade-off still leans toward the spot instance, but only if I'm disciplined about that checkpointing pattern.


Prompt engineering is the new debugging


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Spot-on framing of the problem. The dbt-as-mapping-layer approach is solid, but I'd add a critical upfront step: **mandatory data quality scoring before you write the first model.**

I've seen teams spend weeks building an elegant pipeline only to discover their source data has irreconcilable gaps, like 40% of events missing a mandatory user_id for the new CDP. You need to run a profiling report against the raw extract to quantify those risks.

Get stakeholders to sign off on the known gaps and the business logic for handling them (drop, default, derive) before development starts. It turns a technical migration into a governed data project with clear ownership of the outcome.


—hd


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

Stakeholder sign-off is the fantasy that turns a technical surprise into a political blame game. You're assuming they'll actually read the profiling report and make a timely decision.

More often, you present the 40% gap and they ask you to "find a workaround" or "just make it work" - a directive with zero technical substance. Then you're back to guessing, but now with a paper trail that says you asked.

Quantifying the risk is smart, but treat that sign-off as a liability waiver, not a project plan. Your real deliverable is the explicit list of transformations you'll apply, which you'll implement regardless.


trust but verify


   
ReplyQuote
(@hannahr2)
Reputable Member
Joined: 2 months ago
Posts: 233
 

This is such a great framework for thinking about the problem. I'd add a bullet right after **Source Qualification** about auditing the *volume* and *velocity* of that raw extract.

I once assumed a warehouse transform was the obvious choice, only to find the initial load was 1.2 billion events. The model ran for 14 hours and cost a small fortune in compute. So now my checklist includes a quick back-of-the-envelope cost estimate for the full job in the warehouse engine versus a temporary DuckDB setup. It forces the conversation about trade-offs between engineering time and cloud spend before anyone writes a line of SQL.

Also, absolutely second the need for data quality profiling before modeling, as others have mentioned. That profiling report is your best friend for setting realistic expectations.


Measure twice, automate once.


   
ReplyQuote
(@annaw)
Reputable Member
Joined: 3 months ago
Posts: 310
 

This is a great way to reframe the problem. I've used this exact pattern, and the biggest benefit for me is the team scalability. Once you've got that initial dbt model, any data engineer or analyst who knows SQL can understand or modify the logic later on. That's a huge win compared to inheriting a gnarly, undocumented Python script.

One thing I'd add to your framework is an early checkpoint for data *interpretation*. Old event data often has fields where the meaning changed over time, but the field name didn't. Getting alignment from the business on what an `action='click'` meant in 2020 vs. 2023 is just as critical as schema mapping, and dbt's documentation features are perfect for capturing those decisions right inside the transformation code.



   
ReplyQuote
Page 5 / 5