Skip to content
Step-by-step: Migra...
 
Notifications
Clear all

Step-by-step: Migrating 10 sender domains from SendGrid to AWS SES without downtime.

1 Posts
1 Users
0 Reactions
26 Views
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
Topic starter   [#17708]

Migrating a high-volume sending infrastructure is a pain point most vendors gloss over. They'll sell you on SES's cost savings, but the migration playbook is missing. We just moved 10 domains (≈200M monthly sends) from SendGrid to SES with zero downtime and maintained our reputation. Here's the concrete process, not the slide deck.

**Core Principle:** Dual DNS setup. You run both ESPs in parallel, migrating domain reputation gradually before cutting over IPs and DKIM.

**Phase 1: Pre-flight in AWS SES**
For each sending domain:
1. Verify the domain in SES (both verification and DKIM). SES provides 3 CNAME records for DKIM.
2. Request dedicated IPs (if your volume justifies it) and submit for production access increase. Have your use-case and sending statistics ready.
3. Configure all necessary configuration sets (open/click tracking, custom metrics to CloudWatch).

**Phase 2: Parallel DNS Configuration**
This is the critical step. You do NOT remove SendGrid's DNS records yet.
* Add all 3 SES DKIM CNAME records to your domain's DNS **alongside** the existing SendGrid DKIM record.
* Add a new SPF record that includes *both* ESPs: `v=spf1 include:sendgrid.net include:amazonses.com ~all`
* Maintain SendGrid's existing `s1._domainkey` and `s2._domainkey` records.

Your DNS will look temporarily bloated, but both systems are now authorized to send for your domain.

**Phase 3: Gradual Warm-up & Reputation Migration**
1. Route a small percentage of non-critical traffic (e.g., internal notifications, low-priority alerts) to SES using your application's routing logic. Start at 2-5%.
2. Monitor SES sending quotas, bounce/complaint rates in SES Console and your third-party inbox placement tools (e.g., GlockApps, Netcore).
3. Over 2-3 weeks, gradually increase the traffic percentage to SES. Adjust based on deliverability metrics. The dual DKIM setup means each ESP signs its own mail, and ISPs see consistent signatures from both.

**Phase 4: Full Cutover & SendGrid Cleanup**
Once 100% of traffic is flowing through SES and your reputation metrics are stable (focus on inbox placement, not just opens):
1. Update your application to point all sending endpoints to SES.
2. **Now** remove SendGrid's SPF include and DKIM records from DNS.
3. Update your SPF record to `v=spf1 include:amazonses.com ~all`.
4. Remove the domain from SendGrid's console only after confirming no legacy systems are still attempting to send via it.

**Key Code/Config Snippet - Application Routing Layer:**
You need logic to split traffic. A simple environment variable-driven router:

```python
import os
from typing import Literal

def get_esp_client(for_email_type: str) -> Literal['SES', 'SENDGRID']:
# Control percentage via environment variable e.g., SES_TRAFFIC_PERCENTAGE=30
ses_percentage = int(os.getenv('SES_TRAFFIC_PERCENTAGE', '100'))
# Use a deterministic hash of the recipient email for consistent routing
# This prevents the same recipient getting duplicates during migration
recipient_hash = hash(recipient_email) % 100

if recipient_hash < ses_percentage:
return 'SES'
else:
return 'SENDGRID'
```

**Pitfalls to Avoid:**
* Don't skip the production access request in SES. Your default 200 messages/day limit is a brick wall.
* Monitor for "spoofing" alerts in SendGrid post-cutover. A forgotten service still using old API keys will attempt to send via SendGrid with an SES DKIM signature, which fails.
* SES configuration sets are not retroactive. Set them up before sending real volume.


Show me the query.


   
Quote