Backup and Disaster Recovery

RTO/RPO, AWS Backup, EBS/RDS snapshots, and 4 DR strategies explained with relatable everyday analogies for beginners.

Backup and disaster recovery (DR) is about one central question: if the worst happens, how quickly can we get back to normal?

In everyday terms, it is like making a copy of important documents before storing the originals (backup) and having a fire escape plan that details where to go and what to grab first (DR strategy). Without preparation, a disaster turns into chaos. The same is true for business IT systems.

 

RTO and RPO — The Two Fundamental Metrics of Disaster Recovery

Every disaster recovery plan starts with two questions.

RTO (Recovery Time Objective): "How long can our service be down before it causes unacceptable damage?" If your RTO is 4 hours, you must restore the service within 4 hours of a failure. Think of a restaurant: "Even if there is a fire, we must reopen within 4 hours."

RPO (Recovery Point Objective): "How much data can we afford to lose?" If your RPO is 1 hour, you can tolerate losing up to 1 hour of data. This also means you must back up at least every hour. Think of the restaurant: "We can afford to lose the last hour of order records if necessary."

| Metric | The Question It Answers | Example | |--------|------------------------|---------| | RTO | How fast must we recover? | Service must resume within 4 hours | | RPO | How much data loss is acceptable? | Lose at most 1 hour of data |

Shorter RTO and RPO require more expensive solutions. "Recover in 1 minute with zero data loss" requires the most expensive Active-Active architecture.

 

AWS Backup — One Place to Manage All Your Backups

AWS offers dozens of services: EC2, RDS, DynamoDB, EFS, S3, and more. If you configure backups separately for each service, management becomes complex and error-prone. AWS Backup provides a single, centralized service to manage backups across all of these.

Key Components

A Backup Plan is the policy that defines when backups run, how long they are retained, and when they transition to cheaper storage. For example: "Run a backup every day at 2 AM, keep it for 30 days, then move it to cold storage after 90 days."

A Backup Vault is the logical container where backup data is stored. Think of it as a safety deposit box at a bank.

Cross-Region Backup automatically replicates backups to another AWS region. If a disaster destroys the primary region, you can restore from the backup in a different region.

Cross-Account Backup shares backups with a separate AWS account. Even if the main account is compromised by a hacker or accidentally deleted by an administrator, the backup in the separate account remains safe.

AWS Backup Vault Lock — Protecting Your Backups

Vault Lock makes it impossible to delete or modify backup data for a set period. It protects against malicious insiders and ransomware attacks that try to destroy your backups before encrypting your production data. This is called WORM (Write Once, Read Many) protection. Once Vault Lock is enabled, even administrators cannot disable it.

 

EBS Snapshots — Point-in-Time Photos of Your Disk

A snapshot captures the exact state of an EBS volume at a specific moment in time. Like saving a game, it lets you restore to that exact state later.

Key characteristics:

Incremental snapshots: the first snapshot copies the entire volume. Subsequent snapshots only store the blocks that have changed since the previous snapshot. This saves significant storage space and cost over time.

Fast Snapshot Restore (FSR): normally when you restore a volume from a snapshot, it starts slowly and gradually reaches full performance as data is loaded from S3. FSR pre-warms the volume so it delivers full performance instantly upon restoration. There is an additional cost, but it is valuable when you need to recover quickly during a disaster.

Data Lifecycle Manager (DLM): automates the creation and deletion of snapshots on a schedule. For example: "Create a snapshot every night at midnight. Delete snapshots older than 7 days."

Cross-region copy: copy a snapshot to another region for DR purposes.

Encryption: snapshots of encrypted EBS volumes are automatically encrypted.

 

RDS Backups — Protecting Your Database

Automated Backups vs Manual Snapshots

| Feature | Automated Backups | Manual Snapshots | |---------|-----------------|-----------------| | How created | Automatically every day during a backup window | You trigger them manually | | Retention | 0–35 days (default 7, max 35) | Kept until you delete them | | Point-in-time restore | Yes, to within 5 minutes | Only to the exact snapshot time | | When DB is deleted | Deleted automatically | Retained |

Critical exam point: RDS automated backup retention is maximum 35 days. If regulations require keeping backups for 1 year or 7 years, you must use manual snapshots instead.

!Automated backups versus manual snapshots

Aurora's Special Backup Features

Aurora continuously backs up data to S3 behind the scenes. This enables point-in-time recovery (PITR) down to 1-second precision. Standard RDS supports PITR only to 5-minute precision.

Backtrack is a unique Aurora-only feature. Instead of restoring to a new database instance, it rewinds your current database to a past point in time in-place. This is extremely fast. If someone accidentally runs a DELETE statement that wipes out millions of rows, you can backtrack to 5 minutes before it happened without spinning up a new instance.

 

S3 Versioning and Replication

S3 also provides multiple data protection features.

Versioning: when enabled, every time a file is modified or deleted, the previous version is preserved. If someone accidentally deletes a critical file, you can restore any previous version. The trade-off is increased storage costs.

Cross-Region Replication (CRR): automatically copies objects to a bucket in a different region. Used for DR purposes and for compliance requirements that mandate data copies in multiple geographic locations.

Same-Region Replication (SRR): copies objects to another bucket within the same region. Used to distribute data across development, staging, and production environments, or to aggregate logs from multiple sources into a central bucket.

S3 Object Lock: protects objects using WORM. For a configured retention period, the object cannot be deleted or overwritten. Used to meet regulatory and legal data retention requirements.

 

DR Strategies — Four Ways to Recover from Disaster

There are four main DR strategies, each with a different trade-off between cost and recovery speed.

Backup and Restore

The cheapest but slowest method. You periodically save backups, and when disaster strikes you rebuild the environment from scratch using those backups.

Analogy: after a fire, you collect the insurance money and rebuild the store from the ground up.

RTO: hours. RPO: time since last backup. Cost: lowest.

Pilot Light

Named after the tiny flame that keeps a gas furnace ready to ignite. You keep only the most critical systems (like the database) running at minimal scale in the DR region at all times. Everything else (application servers) is turned off. When disaster strikes, you quickly spin up the rest of the environment around the already-running core systems.

Analogy: the backup kitchen always has the gas on and the pots ready. When disaster hits, you just send in the chefs.

RTO: tens of minutes to a few hours. RPO: minutes. Cost: low.

Warm Standby

A fully functional but scaled-down version of your entire environment runs continuously in the DR region. It handles minimal or no production traffic under normal conditions. When disaster strikes, you scale it up to handle full production traffic.

Analogy: you run a second store at 50% capacity all the time. If the main store burns down, you scale the second store to 100% capacity.

RTO: minutes. RPO: seconds to minutes. Cost: medium.

Active-Active (Multi-Site)

Back to blog list