Skip to content
Notifications
Clear all

Step-by-step: Migrating an RDS instance to Azure with minimal downtime.

1 Posts
1 Users
0 Reactions
24 Views
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
Topic starter   [#12202]

I've been running a mid-sized PostgreSQL database on AWS RDS for a couple of years, but our broader ecosystem is shifting to Azure. The mandate was to migrate with minimal downtime, ideally under 15 minutes. I couldn't find a single guide that covered the full operational and cost checklist, so I reverse-engineered the process and documented the steps.

My primary goal was to avoid the high costs of Azure Database Migration Service for the ongoing sync phase. I used a combination of logical replication and a custom-cutover script. Here's the breakdown of the key phases:

**Phase 1: Pre-migration Setup & Cost Comparison**
* **Source:** AWS RDS PostgreSQL 13.7, db.r5.large (2vCPU, 16GB), 500 GB GP3 storage.
* **Target:** Azure Database for PostgreSQL - Flexible Server, Standard_D2s_v3 (2 vCPU, 8GB). Chose Flexible Server over Single Server for better control and lower compute cost.
* **Cost Snapshot (East US):** The Azure instance came in ~18% cheaper for the equivalent compute tier on a 1-year reserved instance basis. Storage and backup pricing was comparable.

**Phase 2: Establishing Logical Replication**
1. On the RDS instance, you must enable the `rds.logical_replication` parameter and reboot.
2. Create a publication on the source for all tables: `CREATE PUBLICATION mypub FOR ALL TABLES;`
3. On Azure, create the same schema (used `pg_dump --schema-only`).
4. Create a subscription on Azure, pointing to the RDS instance's endpoint. This is where you need to handle the RDS password in the connection string securely.

**Phase 3: The Cutover Script & Downtime Window**
The real trick is the final sync and cutover. You can't just stop writes and wait for replication to catch up—you need to be surgical.
* Scripted steps: Disable app traffic, set RDS to read-only, record the final LSN, wait for the sub to catch up, then promote Azure to read-write.
* The downtime window was 7 minutes and 22 seconds for our 500 GB DB. Most of that was the final sync of large, recently updated tables.

**Phase 4: Post-migration Cost & Performance Validation**
* Validated query performance was within 5% using a replay of sampled production queries.
* The Azure cost breakdown showed savings, but watch the IOPs billing on Flexible Server—it's separate from storage.

Has anyone else tried a similar logical replication move? I'm curious if there are other open-source tools (like pglogical) that could have shaved more time off the cutover.



   
Quote