In the SAA-C03 exam, high availability and disaster recovery (DR) strategies are a critical topic accounting for roughly 20% of the questions. System failures are inevitable. What matters is: "How quickly can you recover, and with how little data loss?" Understanding this is the key to answering DR questions correctly every time.
RPO and RTO — The Two Pillars of Disaster Recovery
Before diving into DR strategies, you need to nail down two fundamental concepts.
Think of it like a personal diary. You write in your diary every single day. Then one day it gets destroyed in a fire.
RPO (Recovery Point Objective) — How much data loss can you tolerate?
If you made a copy of your diary yesterday, you only lose today's entry. If your last copy was a week ago, you lose 7 days of entries. RPO asks: "What is the maximum amount of data we can afford to lose?" The lower your RPO, the more frequently you need to back up your data.
RTO (Recovery Time Objective) — How quickly must you be back online?
After your diary burns, how long does it take to get a new one? If you have to drive to a store and it takes an hour, your RTO is one hour. RTO asks: "How fast must the system be restored after a failure?" The lower your RTO, the more standby infrastructure you need running and ready at all times.
Core principle: The lower your RPO and RTO targets, the more infrastructure you need and the higher your costs.
4 DR Strategies — Think of Them as Insurance Policies
The four disaster recovery strategies map neatly onto insurance policies. As your "premium" (cost) increases, so does your coverage (recovery speed).
Backup and Restore — Basic Home Insurance
The cheapest strategy. You take regular backups, and when a failure happens, you restore everything from scratch.
Insurance analogy: If your house burns down, you file a claim and eventually get money to rebuild. The premium is low, but you might be without a home for months.
RPO: High (you lose all data since the last backup) RTO: High (hours to days) Cost: Lowest Use cases: Non-critical systems, development and test environments
Pilot Light — Always Keep a Minimum Flame Burning
Only the critical core infrastructure — typically the database — runs at minimum capacity at all times. The rest (web servers, application servers) are stopped. When a failure occurs, you quickly spin up the remaining components.
Gas stove analogy: A pilot light on a gas stove is a tiny flame that stays lit all the time. When you need to cook, you just turn the gas knob and the full flame ignites in seconds.
RPO: Medium (DB is continuously replicated, so data loss is minimal) RTO: Medium (tens of minutes) Cost: Low to medium Use cases: Protecting critical databases while keeping overall infrastructure costs down
Warm Standby — A Scaled-Down Full Team Always on Standby
A scaled-down but fully functional version of your entire production environment runs at all times. When a failure hits, you scale it up to full production capacity.
Insurance analogy: You have a small temporary apartment always ready and furnished. It is a bit cramped, but you can move in immediately after a disaster while your main home is being rebuilt.
RPO: Low RTO: Low (minutes) Cost: Medium to high Use cases: Important business systems where some downtime is acceptable but must be minimal
Active-Active — Two Fully Staffed Offices Running Simultaneously
Both AWS regions simultaneously handle full production traffic. If one region goes completely down, the other absorbs all traffic instantly with no interruption.
Insurance analogy: You have two offices in different cities — Seoul and Busan — each fully staffed with identical teams and equipment. If the Seoul office floods, every employee is already working in Busan. Your customers never notice anything happened.
RPO: Near zero RTO: Near zero (seconds) Cost: Highest (running identical infrastructure in two regions simultaneously) Use cases: Finance, healthcare, e-commerce — any system where even one minute of downtime is unacceptable
!4 disaster recovery strategies compared
High Availability Patterns by Service
Understanding how each AWS service implements high availability helps you quickly identify the right answer when the exam asks "which service should you use?"
EC2 High Availability
Distribute instances across multiple Availability Zones (AZs), then combine with an Auto Scaling Group and an Application Load Balancer (ALB).
The essential three-piece pattern: Multi-AZ + Auto Scaling + ALB. If one AZ goes down, traffic automatically shifts to instances in the remaining AZs.
RDS Multi-AZ
RDS Multi-AZ maintains a primary instance and a synchronously replicated standby instance in a different AZ.
Synchronous Replication: every write is immediately applied to the standby before the primary confirms success Automatic Failover: if the primary fails, the standby is automatically promoted within minutes Important distinction: Multi-AZ is for high availability, NOT for read scaling. Use Read Replicas for that.
Aurora Global Database
Aurora maintains 6 data copies across 3 Availability Zones by default. With Aurora Global Database, you can replicate across multiple AWS regions, with secondary regions receiving updates with sub-second latency.
DynamoDB Global Tables
A fully managed multi-region, multi-active database that supports simultaneous reads and writes from multiple regions. This is the textbook example of an active-active pattern.
S3 Durability
S3 provides eleven 9s of durability (99.999999999%). Data is automatically replicated across multiple AZs. You can additionally configure S3 Cross-Region Replication (CRR) to copy objects to a bucket in another region.
Route 53 Routing Policies — Intelligent Traffic Distribution
Route 53 goes beyond simple DNS resolution. It provides intelligent traffic routing that plays a central role in high availability architectures.
| Policy | When to use it | Exam keywords | |--------|---------------|---------------| | Failover | Automatically switch to backup when primary fails | "automatic failover", "DR routing" | | Latency | Route users to the region with the lowest latency | "fastest response", "minimize latency" | | Weighted | Split traffic by percentage (e.g., 90/10) | "canary deployment", "A/B testing" | | Geolocation | Route based on the user's geographic location | "specific country users", "regional compliance" | | Geoproximity | Route based on geographic distance with adjustable bias | "shift more traffic to a specific region" |
Failover routing always works together with Health Checks. Route 53 continuously monitors your primary endpoint. When it fails a health check, traffic is automatically redirected to the secondary endpoint — no human intervention required.
Exam Key Points