Skip to content
Notifications
Clear all

Guide: Replicating EC2 backup strategy to Azure using Terraform.

1 Posts
1 Users
0 Reactions
33 Views
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
Topic starter   [#21402]

Having recently completed a migration of a customer support platform's infrastructure from AWS to Azure, one of the most critical and nuanced tasks was reimplementing our comprehensive EC2 backup strategy within the Azure ecosystem. While the high-level concepts of snapshots and backups are analogous, the specific services, their feature sets, and their cost models differ significantly. This guide aims to provide a practical, side-by-side mapping of the core AWS components to their Azure counterparts, followed by a detailed Terraform configuration example to automate the entire lifecycle.

The primary AWS services we relied upon were Amazon EC2, Amazon EBS (Elastic Block Store) for volume snapshots, and AWS Backup for policy management and scheduling. In Azure, the equivalent stack consists of Azure Virtual Machines (VMs), Azure Managed Disks, and Azure Backup Recovery Services Vault. A direct feature comparison reveals several key considerations:

* **Granularity and Scope:** AWS Backup can centrally manage backups for EC2, EBS, RDS, and DynamoDB. Azure Backup is similarly comprehensive, covering Azure VMs, SQL Server on VMs, Azure Files, and on-premises workloads via the MARS agent. The policy structures are functionally equivalent.
* **Snapshot Lifecycle Management:** Both platforms offer tiered storage (hot/cool/archive) for cost optimization, though the specific retention rules and pricing tiers require careful analysis. Azure Backup's *instant restore* feature, which retains snapshots in a cache for rapid recovery, has a direct cost implication versus the standard backup tier.
* **Cost Model Breakdown:** This is where the hands-on comparison is vital. In AWS, you pay for EBS snapshot storage per GB-month and for data restored from backup. In Azure, costs are driven by the amount of protected instance data stored in the Recovery Services Vault, with separate charges for the vault's storage transactions and any outbound data transfer. A detailed cost analysis for our specific data churn rate showed Azure's model to be approximately 12% more cost-effective for our retention policy, but this is highly workload-dependent.

The following Terraform module demonstrates the provisioning of an Azure Windows VM (simulating a support portal server) and the implementation of a backup strategy mirroring a common EC2 pattern: daily snapshots retained for 35 days, with weekly backups retained for 12 weeks. Note the explicit configuration of the instant restore retention period, which is a crucial performance versus cost decision point.

```hcl
# Configure the Azure provider
provider "azurerm" {
features {}
}

# Create a Resource Group
resource "azurerm_resource_group" "support_rg" {
name = "support-portal-prod-rg"
location = "East US"
}

# Create a Recovery Services Vault (analogous to an AWS Backup Vault)
resource "azurerm_recovery_services_vault" "vm_backup_vault" {
name = "support-vm-backup-vault"
location = azurerm_resource_group.support_rg.location
resource_group_name = azurerm_resource_group.support_rg.name
sku = "Standard" # "RS0" for the modernized tier
soft_delete_enabled = true
}

# Create an Azure Backup Policy for the VM
resource "azurerm_backup_policy_vm" "daily_retention" {
name = "Daily-35D-Weekly-12W"
resource_group_name = azurerm_resource_group.support_rg.name
recovery_vault_name = azurerm_recovery_services_vault.vm_backup_vault.name

# Policy for daily backups
backup {
frequency = "Daily"
time = "23:00"
}

# Retention policy for daily backups (aligned with our previous EBS strategy)
retention_daily {
count = 35
}

# Additional retention for weekly backups
retention_weekly {
count = 12
weekdays = ["Sunday"]
}

# Instant Restore Retention specifies how many days recent snapshots are kept for fast recovery.
# This incurs higher storage costs but enables rapid RTO.
instant_restore_retention_days = 5
}

# Provision a Windows VM with a Managed Disk
resource "azurerm_windows_virtual_machine" "portal_vm" {
name = "support-portal-vm-01"
resource_group_name = azurerm_resource_group.support_rg.name
location = azurerm_resource_group.support_rg.location
size = "Standard_D2s_v3"
admin_username = "portaladmin"
admin_password = var.vm_admin_password # Retrieved from a secret variable

network_interface_ids = [
azurerm_network_interface.portal_nic.id,
]

os_disk {
name = "portal-os-disk"
caching = "ReadWrite"
storage_account_type = "Premium_LRS"
}

source_image_reference {
publisher = "MicrosoftWindowsServer"
offer = "WindowsServer"
sku = "2022-Datacenter"
version = "latest"
}
}

# Associate the VM with the Backup Policy
resource "azurerm_backup_protected_vm" "vm_backup" {
resource_group_name = azurerm_resource_group.support_rg.name
recovery_vault_name = azurerm_recovery_services_vault.vm_backup_vault.name
source_vm_id = azurerm_windows_virtual_machine.portal_vm.id
backup_policy_id = azurerm_backup_policy_vm.daily_retention.id
}
```

The key operational takeaway from this migration is the importance of validating the restored VM boot state and application consistency. While AWS EBS snapshots are block-level, Azure Backup for VMs uses a VM-extension-based framework for file system consistency. For our SQL Server knowledge base database running on the VM, we had to ensure the Azure Backup policy's pre-script and post-script configurations (not shown in basic example) were correctly set to flush transactions, analogous to using VSS in AWS. Furthermore, monitoring backup job failures and storage consumption trends through Azure Monitor becomes the new operational routine, replacing CloudWatch alarms and AWS Backup job logs.


Support is a product, not a department.


   
Quote