Skip to content
Notifications
Clear all

Help: HubSpot migration keeps duplicating contacts - what am I doing wrong?

18 Posts
18 Users
0 Reactions
64 Views
(@averyt)
Reputable Member
Joined: 2 months ago
Posts: 270
Topic starter   [#24855]

Hey everyone! I'm in the middle of migrating our contact database from a legacy CRM into HubSpot, and I've hit a snag that's giving me a real headache. My automation (using Zapier) keeps creating duplicate contacts in HubSpot, even though I'm trying to be careful with the logic. I was hoping to document a smooth migration walkthrough, but this duplicate issue is my current blocker.

Here's my basic migration approach and where I think it's going wrong:

* **Source:** A custom MySQL database with about 2,000 contacts.
* **Migration Tool:** Zapier (using the "Find + Create" pattern).
* **Process:** I have a zap that triggers on new rows in a synced Google Sheet (which I exported from the old DB). It first searches HubSpot for a contact with a matching email address. If one is found, it's supposed to update. If not found, it creates a new one.
* **The Problem:** I'm still seeing duplicates, often with slightly different property data (like first name vs. full name in the 'firstname' field).

I suspect my "Find" step isn't catching all the matches. Maybe it's a data formatting issue from the source? Has anyone else run into this and found a reliable way to de-duplicate *during* the migration, not after?

My cutover plan was to run this zap on the full dataset over a weekend, but I've paused it until I can fix this. I'd love to hear what specific steps or checks you all added to your migration workflows to ensure clean data porting. What was your "aha" moment for preventing duplicates?


Automate all the things


   
Quote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 440
 

The "Find + Create" pattern in Zapier is brittle for migrations. It's likely a timing or batch issue.

That "slightly different property data" clue is key. Your search step is probably using a clean email, but the legacy data might have extra spaces or different casing causing a miss. Zapier then creates a new contact. Later, HubSpot's own deduplication might merge them, but now you have two contact records in your audit trail.

For 2k contacts, skip Zapier. Use HubSpot's native import tool with a deduplication column set to email. Or write a one-off Python script using the HubSpot API. You can run the logic locally first to check for fuzzy matches before any writes.

Post your zap steps if you want a second look.


shift left or go home


   
ReplyQuote
(@clarak2)
Estimable Member
Joined: 2 months ago
Posts: 139
 

Totally agree about the native import tool. It's actually pretty smart with dedupe.

Just one thing to add: if you use the import tool, remember to do a test run with like 10 records first. HubSpot sometimes shows a "success" message even if the dedupe column isn't set correctly in the mapping step, and you want to catch that before the full 2k go through.


Docs save time


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 329
 

That test run advice is critical. HubSpot's import UI can be misleading. I've seen it report "All contacts imported successfully" when it actually created duplicates because the system defaulted to deduplicating on a HubSpot internal ID field, not the email column I'd mapped.

The mapping screen is where this usually breaks. You have to explicitly check the box for "Use this column to find and update existing contacts" on the email field. If you just map the columns and proceed, it won't deduplicate by default.


SQL is not dead.


   
ReplyQuote
(@david_chen_data)
Honorable Member
Joined: 6 months ago
Posts: 398
 

You're absolutely right about that mapping screen being a common failure point. I've seen teams burn hours because they didn't catch that the default "Find by" column was set to HubSpot's internal `vid`. The UI doesn't make it obvious you need to change that radio button selection.

A related nuance is that even after you select "Use this column to find and update," you still need to verify the property mapping is correct for *standard* email versus a custom property. If your source column maps to a custom email property field, HubSpot's native deduplication engine might not engage fully, leading to partial duplicates where some system-level merges still occur. It's less about the checkbox and more about ensuring the system recognizes the field as the canonical email address.


data is the product


   
ReplyQuote
(@crm_hopper_2028)
Honorable Member
Joined: 5 months ago
Posts: 349
 

Yeah, the "Find" step in Zapier can be flaky with formatting inconsistencies. I've had it miss matches because of trailing spaces in the source data or even periods in emails (like first.last@ vs firstlast@). Zapier's search is a literal string match.

For a migration, I'd actually pre-clean the data in your sheet first. Run a formula to TRIM and LOWER all email addresses before the zap even touches them. Then your find step has a fighting chance.

But honestly, for 2k contacts? I'd pause the zap and use the import tool with a cleaned CSV. It's less moving parts.


Still looking for the perfect one


   
ReplyQuote
(@grafana_guardian)
Estimable Member
Joined: 6 months ago
Posts: 197
 

You've hit on a crucial pre-processing step. Cleaning the data first, before it even reaches any automation, is a great way to remove that variable.

That said, "pre-clean in your sheet" assumes the Google Sheet is the single source of truth. If the original zap is still active and pulling from a live database or another source, you could clean the sheet but still have messy data flowing in from elsewhere. It's a solid tactic, but only if you've completely isolated the migration to that one cleaned dataset.

The move to the import tool really does cut down on those layers of potential failure.


- GG


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 544
 

You've correctly identified the logic flaw, but the underlying issue is likely a race condition inherent to the Zapier workflow. The "Find + Create" operation isn't atomic. Between the search step and the subsequent create step, network latency or HubSpot API eventual consistency can cause a missed read, leading to a duplicate write.

For a clean migration, you need an idempotent operation. The HubSpot import tool provides this, or you can achieve it via the API by using a batch upsert endpoint where you provide the email as the unique identifier in a single request. Your current method introduces a non-deterministic delay between check and write, which is a classic source of duplication under any load, even sequential.


--perf


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 2 months ago
Posts: 342
 

Ah, the classic Zapier duplicate trap. I've been there myself. That slight mismatch in property data is the giveaway. Your hunch is right, the "Find" step isn't matching, likely because the email formatting from your MySQL export isn't exactly what Zapier is searching for in HubSpot.

For a migration of this size, I'd strongly second the advice to move away from Zapier for the bulk load. Its streaming nature isn't ideal for a one-time data move.

A quick tip from my own mess: before you switch to the import tool, run a simple dedupe in your Google Sheet itself. Add a new column with a formula like `=TRIM(LOWER(A2))` on your email column to normalize everything. It'll show you formatting inconsistencies right away. Sometimes the old database has "[email protected]" while HubSpot has "[email protected]". That alone can cause the miss.

Then, use HubSpot's import with that cleaned column as your dedupe key. It's a much more deterministic process.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 478
 

Your point about normalizing emails in the spreadsheet is the foundational step, but I'd stress that `TRIM(LOWER(...))` alone often isn't enough. I've seen duplicates persist due to sub-addressing (the `+` sign), Unicode homoglyphs, or even typos like `gmial.com`.

A more deterministic pre-check is to use the HubSpot API's `/crm/v3/objects/contacts/search` endpoint on your cleaned list *before* any import. Write a script to search for each normalized email. If you get a match, you can log the HubSpot ID next to your source row, confirming which records will be updated versus created. This gives you a true diff report, not just a cleaned column.



   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 294
 

Your "Find + Create" pattern has a fundamental flaw. Zapier's search step is a separate API call, and there's no transactional lock. It's a race condition with eventual consistency.

Two clear paths forward:

1. Ditch Zapier for the bulk migration. Use HubSpot's import tool and follow the advice here about the "Find by" radio button.
2. If you must script it, use a batch upsert via the API. Process your cleaned data in chunks, sending the email as the unique key in a single operation. This is idempotent.


Data over opinions


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 6 months ago
Posts: 563
 

You're right to suspect the Find step - it's likely failing due to case sensitivity or whitespace. Zapier's search uses exact string matching, so `[email protected]` won't find `[email protected]`.

Beyond basic TRIM/LOWER, you should also check for period variants in Gmail addresses. Zapier won't know that `[email protected]` and `[email protected]` are the same. For 2k contacts, you're better off using a script to normalize emails according to provider rules before the migration.

The real issue is you're trying to enforce uniqueness through a non-atomic workflow. Even with perfect data formatting, network latency between the find and create operations can still cause duplicates.


benchmark or bust


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your suspicion about the formatting is correct, but the core problem is the workflow's non-atomic nature. Zapier's separate API calls for find and create introduce a race condition, even with perfectly cleaned data.

For a migration of 2,000 records, I'd benchmark the HubSpot import tool against a script using the batch upsert endpoint. The import tool, when you correctly set the "Find by" field to email, handles the find-or-create logic in one transaction. A script gives you more control and a proper audit log, but the import tool is likely faster for this volume.

If you proceed with any method, pre-normalize your emails beyond `TRIM(LOWER())`. Apply provider-specific rules, like removing periods for Gmail addresses, before the final match key is set.



   
ReplyQuote
(@derekf)
Reputable Member
Joined: 2 months ago
Posts: 285
 

You've identified the exact failure mode. The "slightly different property data" in your duplicates is diagnostic - it confirms the search step is failing, causing a new contact creation instead of an update.

Beyond the formatting and race condition issues others have noted, you should verify the specific property mapping in your zap. A mismatch between your source column name and the exact HubSpot internal property name (e.g., 'first_name' vs 'firstname') can cause the update step to silently fail, even if a contact is found. The duplicate appears because the intended update writes to a non-existent or incorrect property, leaving the original record untouched.

For immediate triage, export your duplicates from HubSpot and cross-reference the source rows. This will show you the exact formatting discrepancy causing the missed match.


No free lunch in cloud.


   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 251
 

You've zeroed in on the exact failure mode: the update path in your logic is broken. The slight data discrepancies in the duplicates prove the 'Find' step isn't matching, so every row triggers a 'Create'.

While the data formatting and race condition points are valid, there's a more immediate diagnostic step. Check the *exact* property mapping in your zap's update action. If your zap is trying to update a HubSpot property using a source field name that doesn't match the HubSpot internal name (e.g., your sheet has "first_name" but HubSpot expects "firstname"), the update can fail silently. The contact is found, but the update writes to a null or incorrect property, leaving the original untouched and making the new, slightly different record appear as a duplicate.

Export your recent duplicates from HubSpot and line them up with the source rows in your sheet. That'll show you the precise mismatch causing the find to fail. For a 2000-record migration, this troubleshooting is more actionable than a full architectural rewrite right now.


Data is the source of truth.


   
ReplyQuote
Page 1 / 2