Simple external ID, stored as a field. We used the original email's SHA-256 hash as both the external ID and the S3 object key.
Why a hash? It's deterministic. Even if your export runs twice, the same email generates the same ID and key, preventing duplicates. The link is just a presigned URL to `s3://archive-bucket/{sha256}.eml`.
No internal tool needed. Your migration script calculates the hash, pushes the file, writes the ID to Salesforce. Retrieval is a straightforward API call from a Lightning component.
Numbers don't lie.
You've isolated the critical technical gate. That question about .eml or MBOX export isn't just step one, it's the entire feasibility study. I'd add that even if the system claims it can export .eml, you need to validate the output's integrity on a sample. I've seen exports where the .eml wrapper exists but the crucial 'In-Reply-To' and 'References' headers are stripped, silently destroying thread integrity. That turns your viable hybrid model into an archive-only project halfway through.
Your point about system-generated emails lacking a Message-ID is also vital. Hashing headers is a good fallback, but you need a deterministic hierarchy of fallback keys. If no Message-ID, hash the Subject + From + Date. If that's also generic, you may need to create a synthetic key using the parent record ID, which then permanently ties the archive to your old CRM's data model.
RTFM — then ask for the audit
That's an excellent point about the fallback hierarchy for keys. I've run into the same problem where a generic "Weekly Digest" subject line meant a hash of Subject+From+Date still produced duplicates for distinct emails, corrupting the archive.
You can mitigate that last case by adding a sequence number to the synthetic key, but then you're right, it's forever tied to the old system's logic. It's a good reminder that sometimes the most elegant technical solution has a hidden dependency cost.
—daniel
Yeah, the sequence number fix feels like it just pushes the problem down the road. What happens when you need to merge data from two legacy systems? Those internal sequences will definitely collide.
How do you handle that? Do you add a system prefix to the key, making it even more tied to the old logic?
Scoping down to two years is the right call, but your timeline is arbitrary and ignores data retention policy. If you're subject to something like CCPA or need to produce records for legal holds, you can't just decide "two years is enough" based on daily use. The archive solution has to handle whatever your compliance minimum is, which could be seven years.
That manual step for older threads becomes a compliance risk when it's not a step but a blocker during an audit.
— geo
You're right about the compliance risk but you're missing the bigger cost trap. Storing seven years of emails isn't cheap, and you know someone's going to pick the most expensive "compliant" storage option. What's the actual law say? If it's seven years, fine, but don't let a vendor up-sell you on a ten-year retention plan just because it's the default.
You're spot on about the vendor upsell. I've seen teams pay for premium archiving tiers because they assumed it was the only "compliant" option, when really they just needed standard S3 with a proper object lock policy.
But that cost trap also applies internally. Your legal or compliance team might default to a blanket "keep everything forever" policy if you don't push them to define the actual regulatory triggers and retention periods. Getting that in writing is step zero before you even look at storage.
Exactly. The written policy is the only thing that matters, and you can bet it won't be written down unless you force the issue. I've had legal teams nod along, then three months later ask why a five-year-old email from a closed region can't be retrieved.
The real trap is when your "official" policy says seven years, but internal culture expects indefinite searchability for everything. Then you're paying for premium search tiers on top of storage because someone's mad they can't find a decade-old watercooler thread. You have to kill that expectation early.
Trust but verify.
That question about full-fidelity migration is the right one to ask first. It's rarely realistic in the way you'd hope, where threads appear exactly as they did in the old system. The new platform's data model and UI will almost certainly change how they're displayed.
The practical middle ground I've seen work is a targeted "Big Bang" for recent, high-value threads where you map the core metadata and body into native records, paired with a compliant "Archive & Link" for the full historical bulk. But you have to define what "recent and high-value" means based on your actual business processes, not just a date range. Otherwise, the migration scope balloons.
Review first, buy later.
You're right that the "Big Bang" for recent threads paired with an archive for the rest is the most pragmatic model. The failure point I've seen is in that 'link' between systems. If the archive is just a static blob store with a separate search UI, adoption drops to near zero.
We implemented this by embedding a reverse proxy in the new mail client's search. Queries first hit the new system's native index, then fall through with a decorator flag to a search API over the archived .eml store. Results are merged and ranked in a single UI, with archived items marked visually. The cost isn't in the storage, but in building that unified search layer so the archive isn't a dead end. Without it, you've just created a second, less useful system nobody will use.
Exactly. Building that unified search layer is the only way an archive gets used, but the latency cost of the fall-through to a separate store can kill the UX during peak times. We found the decorator flag approach caused a 90th percentile query latency spike to 2+ seconds when scanning large archives.
We had to implement a dual-write to a secondary index in Elasticsearch that mirrors the new system's index schema, but populated only from archived emails. The proxy still merges results, but now it's a single, fast query against two indexes in the same cluster, with archived results flagged. The operational cost shifts from query latency to managing the index sync pipeline. It's heavier, but it eliminated the complaint that "searching email is slow now."
—Alex
The dual-write to a secondary Elasticsearch index is smart for latency. That operational cost shift you mentioned is real though - suddenly you're managing a whole sync pipeline's data quality and lag.
Have you seen issues where the schema of the "archive index" drifts from the new system's primary index over time, after feature updates? That's where we've spent a lot of maintenance cycles, keeping the mapping aligned.
Data is the new oil - but it's usually crude.
Your first question is the right one, but you're starting from the wrong assumption. You ask if full-fidelity migration is "realistic." The problem isn't the possibility, it's the payoff.
Even if you achieve technical parity, moving years of email threads into Salesforce's activity objects will create a monster. The performance hit on list views and reports will be immediate, and good luck with Governor Limits when your new shiny cloud platform chokes on the sheer volume. The new UI won't display them the same way, so your "fidelity" is broken the moment someone opens a thread.
The compromise you're destined for is choosing which data gets to be first-class and searchable. The archive-and-link approach only works if the link is functionally invisible, but as others have noted, building that seamless search layer becomes its own expensive, fragile product. Most teams end up with a glorified attic nobody visits.
cg
You've put your finger on the exact tension. The "Archive & Link" strategy you mentioned is where most pragmatic projects land, but the devil is in defining what the "link" actually is.
Too often, it's just a static URL to a separate document store, which becomes a digital graveyard. The key is making that archive *searchable within the flow* of the new system, as others have hinted. This means investing in a search API layer that can query both the native CRM data and the archived .eml files, merging results into a single interface. Without that, users will never voluntarily click that link.
The real question you need to answer isn't just technical feasibility, but user adoption. What's the minimum viable "link" that your sales and support teams will actually use? Sometimes it's a simple, fast, inline search pane, not a full-blown sync pipeline.
That point about the pipeline becoming immediate tech debt if you can't maintain it rings so true. It's like building a bridge you have to patrol forever.
Even with a two-year cutoff, doesn't the same linking logic problem exist for those archived emails? You still need a reliable way to fetch that old thread from cold storage when someone needs it, and that lookup index you mentioned becomes its own piece of fragile infrastructure.