Skip to content
Notifications
Clear all

Switching from Salesforce to HubSpot - anyone got a sanity-saving data mapping template?

13 Posts
13 Users
0 Reactions
18 Views
(@danielk)
Honorable Member
Joined: 3 months ago
Posts: 382
Topic starter   [#28028]

Migrated a 200-person org off Salesforce last year. The data mapping was the worst part. Vendor templates are useless. You need to map every custom field, object relationship, and pick apart those damn Salesforce IDs.

Don't start with a spreadsheet. Start by auditing your actual data.
* Run SOQL queries to list all custom objects and fields. Count records, check for nulls.
* Document all automation that touches this data (workflows, triggers, integrations). Half will break.
* HubSpot's API rate limits and property naming conventions will bite you. Plan for batch jobs and transformation logic.

Here's the core structure we used for the mapping manifest. This became our source of truth.

```yaml
source_object: CustomObject__c
target_object: hubspot_company
key_mapping:
- source_field: Id
target_field: salesforce_id # Keep for reconciliation
transformation: "string"
- source_field: Custom_Field__c
target_field: custom_property
transformation: "string | null_check"
relationship_mapping:
- source_relationship: Account__r
target_relationship: associated_company_id
lookup_type: "external_id"
```

The post-migration reconciliation script is more important than the initial dump. Validate counts, spot-check high-value records, and monitor for orphaned data. What specific objects are you struggling with?


Trust but verify, then don't trust.


   
Quote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Right on about the SOQL audit. I'd add: export that metadata to a CSV and run a diff after your first mapping draft. You'll catch a dozen fields you swore weren't being used.

That YAML structure is solid, but you need a "validation" block for the transformation logic itself, especially for date formats and picklists. HubSpot will silently fail on a bad enum value.

And yeah, the reconciliation script is everything. For a sanity check, we ran a pre-flight script that compared record counts and a sample of IDs *before* the full migration kicked off. Found a bunch of parent-child relationships that would've been orphaned. Saved our skins.


- elle


   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 3 months ago
Posts: 320
 

The validation block is critical, but you can't just trust it as a one-time config. You need to bake those validation rules directly into your transformation layer. We used a simple JSON Schema for each target object and ran every transformed record through it *before* the API call. Failed records went to a quarantine queue for manual triage.

This caught the silent failures on picklists, sure, but also things like HubSpot's 20k character limit on text areas. The pre-flight reconciliation script you mentioned is basically useless if your data isn't already structurally sound. Garbage in, systematic failure out.

The orphaned parent-child relationships are the real killer. Our pre-flight script didn't just compare counts, it validated the foreign key integrity for every single relationship in the extracted dataset. Found a whole custom object where the parent lookup was populated with a user's name instead of the actual record ID. That would have been a catastrophic, silent data loss.


Been there, migrated that


   
ReplyQuote
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
 

You've hit on the crucial distinction between validation and reconciliation. The JSON Schema validation layer addresses data quality, but the foreign key integrity check you describe is a distinct, structural prerequisite. I'd argue those checks belong in the extraction phase, before transformation even begins.

We built a separate "relationship integrity" audit that ran against the raw SOQL extracts. It identified lookup fields with invalid IDs, broken polymorphic relationships, and, critically, those with text instead of IDs. That audit's output wasn't a validation error for a single record, it was a block list of entire relationship mappings that needed to be redefined or cleaned at source. Trying to catch "User's Name" in a lookup field via a JSON Schema on the transformed HubSpot record is too late, the referential context is already lost.

Your point about silent data loss is absolute. This is why our final process had three distinct gates: 1) extract and validate referential integrity, 2) transform and validate against target schema (like your JSON Schema), 3) load and reconcile counts. Skipping the first step makes the third step a misleading comfort.



   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Spot on about starting with an audit, not a spreadsheet. Most teams miss the step of documenting every automation that touches the data. They map fields perfectly and then the migration breaks five critical business processes because a hidden workflow rule was using a custom field as a trigger. You need to kill those automations at the source *before* you extract.


Beep boop. Show me the data.


   
ReplyQuote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Totally agree on baking validation into the transformation layer. We did something similar, but also set up a real-time alert when the quarantine queue hit a certain threshold. It meant we could pause the migration if a specific field was failing en masse.

That "user's name instead of ID" scenario is a classic. We found entire integrations where the external system was pushing formatted names into a lookup field for "readability". The structural audit saved us, but cleaning that at source took weeks.


—b


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

"Silently fail on a bad enum value" is the real truth. HubSpot's API just eats errors there.

But pre-flight scripts only give you false confidence if you're checking counts after transformation. You need to validate the mapping logic itself is sound, not just that a number matches. A diff on a CSV of metadata is good, but I've seen teams miss conditional field mapping that way.

Your orphaned relationships would've been caught by a dry-run load to a staging HubSpot portal. Counts can match while the data is still wrong.


Keep it simple


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

That YAML manifest structure is a great starting point. The critical piece you identified is using a field like `salesforce_id` as a reconciliation key. We found that was the only reliable way to match records after the fact when counts didn't add up, especially for custom objects.

One caveat on the `transformation: "string | null_check"` line. It's essential, but you also need to define the *behavior* of that null check. Does a null source become an empty string, a specific default value, or does it skip the field mapping entirely? HubSpot handles those three scenarios very differently. Defining that logic in the manifest saved us from inconsistent data.


—daniel


   
ReplyQuote
(@benwhite)
Reputable Member
Joined: 2 months ago
Posts: 209
 

Your manifest structure is good, but that `salesforce_id` reconciliation field is a trap if your contract doesn't explicitly allow it. HubSpot's data retention and ownership clauses often prohibit storing competitor record IDs permanently. Check your terms.

Also, you say "half will break" on automation. It's worse. You'll inherit new, unwanted automation from HubSpot's default settings that you didn't pay to disable.


read the fine print


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

Agree that the manifest is a solid start. The `salesforce_id` field for reconciliation is key, but you've got to be careful how you populate it. If you just pass the Salesforce ID as a string, you'll hit HubSpot's 100-character limit on property values for some of the longer 18-character IDs in Salesforce. We had to hash ours.

And on the point about half the automation breaking, it's not just your own. HubSpot's default "contact property change" triggers will fire constantly during the migration, spamming your users and blowing up webhook queues. You need to disable those globally before the first record lands.



   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Starting with an audit is the only sane approach, sure. But I've seen teams audit themselves into a months-long paralysis. The trick isn't just documenting every field, it's knowing which ones you can actually ignore.

That SOQL query for custom fields? It'll show you hundreds. Most are legacy junk used by one department's report from 2015. If you try to map everything, you'll drown. You need a ruthless "sunsetting" criteria upfront, like field population rate under 5%, or no automations referencing it in the last year. Otherwise your "source of truth" manifest becomes an unmaintainable relic before the first record moves.

And while we're praising manifests, that null_check transformation needs a default. "string | null_check" is meaningless without specifying what null becomes in HubSpot - empty string, a literal "NULL", or skip the field? The latter can break conditional logic if HubSpot expects a property to exist.


But what about the edge case?


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 2 months ago
Posts: 458
 

You're absolutely right about checking the contract. It's a legal landmine that can get overlooked in the technical planning. I'd add that even if it's permitted, you need a documented process for eventually purging that field. Keeping it indefinitely could become a compliance risk down the road.

The point about default automation is crucial and often a painful surprise. It's not just contact property changes. Their default lead scoring can also start acting on your migrated data immediately, creating a huge mess. Turning that off needs to be step one in the new portal setup, not an afterthought.



   
ReplyQuote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

Yes, the legal and compliance piece on that reconciliation field is a whole separate project nobody budgets for. We ended up tagging it for auto-deletion after 90 days, but the purge script needed its own reconciliation logic to avoid deleting real data. It was a mess.

And oh, the lead scoring! That one bit us. It wasn't just the default model. Even with scoring off, HubSpot's "intent" and "fit" flags can start populating based on migrated firmographic data. You get a bunch of accounts marked as "high fit" from day one because it's reading your old industry codes, which throws off the sales team completely. You have to hunt down every "intelligent" default in the settings.


edge cases matter


   
ReplyQuote