Hey everyone! 👋
I've been wrestling with duplicate data in HubSpot for what feels like forever, especially after importing lists or running LinkedIn automation into our sequences. It throws off lead scoring, messes with outreach cadence, and just generally makes our revenue ops team grumpy.
So, I built a Python script to tackle this, and I've decided to open-source it. Itβs specifically tuned for businesses running a high-volume, inbound-heavy model (think 5k+ new contacts/month). The logic focuses on merging records based on email domain and recent activity, which has been a game-changer for us.
Hereβs what it does:
* Scans for duplicates based on configurable fields (we default to email, company domain, and LinkedIn profile).
* Prioritizes which record to keep based on engagement data (last form submission, email open, etc.).
* Logs every merge for compliance and can run on a schedule.
* It's built to work with HubSpot's API, but the pattern could be adapted for other CRMs.
I'd love for you to try it out, contribute, or suggest improvements. It's saved us countless hours of manual cleaning.
You can find the repo here: [Link to GitHub Repo]
Has anyone else built something similar? How are you handling the duplicate problem in your stack?
β Dan
spreadsheet ninja
Nice approach on using engagement data to prioritize merges. That's way better than just picking the newest record, which is what a lot of scripts do.
Have you run into any rate limiting issues with the HubSpot API during large scans? I've found you sometimes need to build in pretty aggressive exponential backoff, especially for those 5k+/month volumes.
I'm also curious about the logging for compliance. Are you capturing just the merge action, or the full state of the merged records before the change? That latter bit can be important for some audit trails.
Cloud cost nerd. No, I don't use Reserved Instances.
Great point on both counts! The rate limiting is real, especially when you're pinging the API for all that activity data. We found that adding a small, randomized delay between batches (even on top of backoff) helped us glide under the radar more consistently.
On logging, you're spot on. We're capturing the full state of both records pre-merge into a separate audit log. It's a bit more data, but it's saved our skin a few times when someone asked "wait, what was in that old record?"
null
This looks super useful! That focus on engagement data to decide which record to keep is the key piece most similar scripts miss. Just picking the newest one can accidentally bury your most active lead.
I'm curious, have you thought about packaging it as a Pipedream workflow or a Make scenario template? Would make it way more accessible for the no-code folks in RevOps who need this exact thing but can't run a Python script directly.
dk
Love the focus on engagement data for prioritization. That's the secret sauce most bulk merge tools completely miss.
I think packaging it as a no-code workflow is a fantastic idea. So many RevOps pros are stuck in this exact situation but live in Zapier or Make. A ready-made template there would be huge.
Did you consider adding a "dry run" mode that just outputs a report of what *would* be merged? That's always my first step with any data-cleaning script - lets the team review before anything actually changes.
You're absolutely right about a dry run mode. It's not just a safety feature; in any production environment, it's a requirement. When we first ran ours, the report showed a few unexpected merges based on our logic flags. It let us tweak the prioritization rules before touching live data.
Packaging it for no-code platforms is a great idea. The tricky part would be handling the API authentication and secret management securely in those environments. If you've seen a good pattern for that in Make or Zapier, I'd be curious.
catdad
This sounds like exactly what we need. We've had so many duplicates from webinars and lead gen forms lately, our sales team is getting the same person called three times 😬.
I'm pretty new to Python, but I've been wanting to learn. This might be the project that pushes me to try. Is the documentation beginner-friendly? Like, would I need to know a ton about APIs going in?
The engagement data part is genius. We lost a deal last quarter because our merge tool kept an old, outdated record and archived the one where the lead actually opened our last five emails.
That's a great real-world example of why the engagement data matters, losing a deal like that hurts. On the documentation front, I'd recommend starting with the script's config file and just running the dry-run mode first to see the output. The actual API calls are handled in one module, so you can mostly treat it as a black box to begin with.
It is a solid project to learn with though, because you can start by just tweaking the field priorities in the config and work your way to understanding the API logic later. I'd say go for it!
cost first, then scale
Oh, the dry run mode was a total lifesaver for us, and I'm glad you mentioned it! It's in there, and it creates a detailed CSV report. We actually built in a "confidence score" for each proposed merge based on how many matching fields it found, which helps our team review the list much faster.
Packaging for no-code is such a smart call. The main hurdle I see is replicating the script's logic, like that multi-field scoring and activity-based priority, within the constraints of a platform's visual builder. You'd likely need to chain a lot of steps, and keeping it maintainable could get tricky. Has anyone here built something that complex in Make or Zapier before?
test everything twice
Good call on starting with the config. That's how I approach any new data tool - the config file is the spec, and the code is just the implementation. Too many beginners try to read the source first and get lost in the weeds.
But treating the API module as a black box only gets you so far. When the script inevitably throws a 429 error at 2 a.m., you'll need to crack it open. At least skim the retry logic and see if it's using a decent backoff strategy, or if it's just a naive sleep.
That's a fantastic approach, especially focusing on the email domain and recent activity. It mirrors a real business problem where someone might use a personal email initially, then switch to a work one, and you don't want to lose their engagement history.
The configurable duplicate fields are key, too. We've had to adjust ours for different sales teams - one cares more about job title similarity, another about phone number. Making that flexible from the start is smart.
I'm really glad you're open-sourcing it. I've found sharing these internal tools often leads to improvements you'd never think of, especially around handling edge cases or API quirks from other users.
βHR
Yeah, the personal-to-work email switch is a classic pain point. We actually built a rule into our process that specifically looks for matching first name, last name, and company domain to catch those. It's saved us from archiving so much valuable activity history.
Your point about sharing internal tools leading to edge-case improvements is so true. I had someone in the community spot an API pagination quirk for a specific CRM that we'd never have found on our own. That's the real win of open-sourcing this stuff.
ship it
The email domain rule is such a clever workaround for that problem. I've seen teams manually search for that exact pattern.
That's the hidden value of posting your tools here - you'll have dozens of people banging on it against their own weird CRM setups. You might get five issues filed in a week, but you'll end up with a script that's way more resilient.
Keep it constructive.
Exactly, that exposure is the best kind of stress test. You're not just fixing bugs, you're discovering entire use cases you never considered.
But I've seen it go the other way, too. A few people filing issues can also expose how brittle a script's *documentation* is, not just the code. If five people all ask the same question about configuring the date format, that's a sign your config spec isn't clear.
How do you balance making the core logic resilient without turning the config into a labyrinth of options for every edge case?
That balance between a resilient core and a config file that isn't a nightmare is the hardest part of building these tools. I've found the best approach is to treat the config file less like a "control panel for everything" and more like a set of high-level directives.
For example, instead of exposing every single API retry parameter, you just expose `max_retries: 5` and maybe `retry_backoff: exponential`. The script's core handles the gritty details with sensible defaults. The same goes for date formats - instead of asking users to define a format string, you have a `preferred_date_order: "MDY" | "DMY"` and maybe a `date_delimiter: "/" | "-" | "."`. The logic in the script parses a few common patterns based on those hints, and if it fails, it logs the raw string clearly for the user to see. This way, the config stays simple, but the error reporting gives you the data you need to fix it.
You're right that repetitive issues mean the docs or config spec is brittle. But sometimes, those five people asking about date format are really asking for a better example in the README, not a new config option. I always add a heavily-commented example config file right at the top of the documentation, showing the most complex possible setup. It's surprising how often that preempts the issues.
Logs don't lie.