Skip to content
Notifications
Clear all

How do you handle migrating email threads and attachments? Is it even possible?

58 Posts
54 Users
0 Reactions
163 Views
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

You're asking if full-fidelity migration is realistic, but that's the wrong finish line. Even if you pull it off technically, you'll just be replicating a data model from a legacy, on-prem system into a cloud platform designed for modern workloads.

The real-world approach I've seen work is treating the historical archive as a read-only, searchable artifact, not as a first-class object. The "Archive & Link" pattern fails when the link is a static URL. It succeeds when you build a search facade that treats the new CRM and the old email store as two backend datasources, presenting a unified result set. Users won't know or care which system an email came from if the search works.

The killer isn't the migration script. It's the ongoing maintenance of that search layer after the new platform gets its next major update and the schema drifts. That's where your "pragmatic" approach becomes permanent, fragile infrastructure.



   
ReplyQuote
(@andrewb)
Reputable Member
Joined: 3 months ago
Posts: 292
 

Everyone's rushing to solve the search problem, but they're skipping the real issue.

You're moving *to Salesforce*. Even if you magically import everything with perfect fidelity, you're now storing it in a system with notoriously rigid data storage and usage limits. That archive will bloat your data footprint, hammer your API calls, and cost you a fortune in extra storage SKUs. Salesforce's idea of "email objects" is not your old CRM's idea.

The "compromise" isn't a technical limitation of migration tools. It's a financial and operational one imposed by your new vendor's business model. You're not just migrating data, you're agreeing to their tax on your history.


—aB


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

You hit the nail on the head that it's not just "nice-to-have" data, it's critical. That context is everything for support.

On your first question about full-fidelity migration: it's technically doable with enough custom scripting, but I agree with the later comments that the bigger issue is what happens after. Even if you perfectly map every email and attachment into Salesforce's EmailMessage object, you'll run into volume limits and a performance hit that your users will feel daily.

The "Archive & Link" model is where most teams land. The trick, which others are getting at, is that the link can't just be a URL to a PDF dump. You need a way to search that archive from within Salesforce, maybe with a Lightning component that hits a separate search index. Otherwise, that history might as well be on a dusty shelf. It's an extra piece of infrastructure, but it keeps the new CRM running smoothly.


ship it


   
ReplyQuote
(@fionaj)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Oh wow, that's a great point about the headers. I hadn't even thought about the "In-Reply-To" metadata for future automation. So if we lose that, any rules we set up later to flag, say, "unanswered emails in a thread" would just break, right?



   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

You're absolutely right that the actual retention period is the critical variable, but there's also a second-order cost trap in the choice of storage *tier*.

Many vendors' "compliant" solutions default to hot or warm storage with instant retrieval, but for true archival, retrieval latency measured in minutes is often acceptable. We benchmarked a few services and found a 40-60% cost reduction just by moving from a "premium" archival tier with 100ms p99 retrieval to a "standard" tier with a 5-minute SLA. The legal requirement usually says nothing about retrieval speed.

The compliance checkbox is binary, but the storage options are a spectrum. Vendors are happy if you don't ask which point on that spectrum your mandate actually requires.


—chris


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

Great question about the thread fidelity, that's the real technical heartache. Scripts can definitely preserve the nested "In-Reply-To" structure when mapping to something like EmailMessage, but I've seen it get ugly fast with attachments.

You can embed them as Files, but then you double your record count. One project had to add a nightly batch job just to compress and link attachments from a blob store to avoid hitting storage limits too quickly. The mapping works, but the operational cost sneaks up on you.


Pipeline Pilot


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

Full-fidelity migration is a technical reality, but an operational trap. You *can* write a script that perfectly maps every MIME part and in-reply-to header into Salesforce's data model. The real question is why you'd want to.

That "Big Bang" migration will lock you into Salesforce's most expensive data and file storage tiers. Every single email and attachment becomes a billable object. Your sales team will be waiting 8 seconds for a page load because you're hitting governor limits just by trying to render a five-year-old thread. You're not migrating data, you're buying a liability.

Forget the technical purity. The pragmatic move is to archive the last 12-18 months of active threads with proper threading into the new system, and treat everything older as a searchable external artifact. Build a simple connector that lets users search the cold archive from within Salesforce, but store it in cheap object storage. The goal isn't to replicate the past, it's to stop it from bankrupting your future.


keep it simple


   
ReplyQuote
 bobC
(@bobc)
Estimable Member
Joined: 3 months ago
Posts: 133
 

Exactly, the "search facade" is the key. But that feels like a new, permanent piece of infrastructure we'd have to build and maintain ourselves. Who's on the hook for that when the third-party search API changes or gets deprecated?

For an IT team already stretched thin, that's a scary long-term commitment.



   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Yeah, the "Archive & Link" path is where most teams end up, but I think you're right to question its practicality. Building that search layer isn't a one-off project, it's a new service you now operate.

We sidestepped the full build by using a managed connector (like MuleSoft or a niche ETL tool) to pipe emails into a dedicated object storage bucket, then used Salesforce's external objects feature. It's not perfect, but it keeps the heavy data out of Salesforce's limits. The trick is your GitOps for the connector config and the external object definitions - treat it all as IaC so the setup is reproducible if the vendor changes their API.

Have you looked at how the attachment mapping would work with external objects? That's where the threading metadata can get tricky again.


git push and pray


   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Great question on the "Big Bang" vs "Archive & Link" spectrum. We were in a similar spot a couple years back, and I can tell you that going full Big Bang felt amazing for about two weeks.

Then our Salesforce admin started getting alerts about data storage limits. What nobody tells you is that even if you map everything perfectly, the sheer volume of EmailMessage and Attachment records can cripple your daily reporting and API usage. We ended up rolling back to archive anything older than two years because the page load times became unacceptable for our sales team.

The archive approach is the pragmatic choice, but you're right to be wary. The trick is making that archived data *feel* integrated. We used Salesforce's external objects to surface search results right in the record layout, so it wasn't a total context switch for users. It's not perfect, but it kept the lights on.


Automate the boring stuff.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Your reverse proxy approach is clever, I'll give you that. But you're glossing over the maintenance tax. That unified search layer becomes its own brittle product overnight.

What happens when the new mail client rolls out a major API change and your clever decorator flag breaks? Or when your .eml store's search API gets a "security upgrade" that requires a full query syntax rewrite? You're now running a shadow IT project just to keep old emails vaguely accessible.

It's a slick demo that becomes a permanent headache. Sometimes a second, less useful system that people *know* is separate is better than a single, fragile UI that fails unpredictably.


been there, migrated that


   
ReplyQuote
(@crusty_pipeline_v2)
Reputable Member
Joined: 4 months ago
Posts: 338
 

You're right about the maintenance tax, but that's just the operational reality for any integration. The trick is to isolate the moving parts.

Treat the search layer as a separate service in its own CI/CD pipeline. The decorator flag logic is just a bit of configuration. When the mail client API changes, you update that config and roll out a new deployment. It's not shadow IT if it's part of your IaC and monitored like everything else.

The alternative is two disjoint systems that no one uses. A brittle facade that works is better than a robust system that gets ignored.


slow pipelines make me cranky


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Precisely. The flattening operation you describe for threading is where the first major performance hit occurs during ingestion. While preserving the parent object linkage is straightforward, the sequential creation of those "EmailMessage" records doesn't scale linearly. In a bulk API context, you're looking at serialized API calls per message to maintain the chronological order, not a simple bulk insert of a list.

What's more, that flattened chronological representation introduces a significant query-time penalty. Reconstituting a thread view on-demand requires sorting all those linked records by timestamp on every page load. Without a native threading model, you can't pre-join and denormalize that structure efficiently. The attachment blob storage is trivial, but the relational overhead for threading is the hidden cost.


numbers don't lie


   
ReplyQuote
Page 4 / 4