Disaster Recovery (DR) and cost visibility are the final two tasks of SAP-C02 Domain 1. While they have fewer questions, DR and cost concepts appear repeatedly across other domains (D2, D3), making them essential to master completely.
The core of DR design is defining "how much downtime and data loss the business can tolerate" first, then selecting the strategy that fits. Cost visibility is about "understanding where money is being spent and finding optimization opportunities."
Understanding RTO and RPO
RTO (Recovery Time Objective) is the maximum acceptable time from failure to service restoration. RPO (Recovery Point Objective) is the maximum acceptable data loss measured in time.
For example, if RTO is 1 hour and RPO is 15 minutes, the service must be restored within 1 hour of a failure, and at most 15 minutes of data can be lost.
Shorter RTO and RPO require more infrastructure and higher cost. The exam always tests this trade-off.
The 4 DR Strategies
AWS defines 4 DR strategies where cost and recovery speed are inversely related.
| Strategy | RTO | RPO | Cost | Description | |----------|-----|-----|------|-------------| | Backup and Restore | Hours | Hours | Lowest | Restore from backups. Provision infrastructure from scratch | | Pilot Light | Tens of minutes | Minutes | Low | Only core infrastructure (DB) runs continuously. Rest provisioned on failure | | Warm Standby | Minutes | Seconds to minutes | Medium | Scaled-down full environment runs continuously. Scale up on failure | | Multi-Site Active-Active | Near zero | Near zero | Highest | Both regions serve traffic simultaneously. Instant failover |
In Pilot Light, only database replicas (such as RDS Cross-Region Read Replicas) are maintained in the DR region. Web and app servers are launched from AMIs when a failure occurs. Warm Standby runs all tiers in a scaled-down state, so only scaling up is needed.
Elastic Disaster Recovery (DRS) is an automated DR service for rehost scenarios. It continuously replicates on-premises or other cloud servers to AWS and can recover them as EC2 instances within minutes of a failure.
Cost Visibility Tools
Understanding and optimizing costs at the organizational level is a key Solutions Architect responsibility.
Cost Explorer visualizes costs and usage. It analyzes costs by service, account, and tag, and includes forecasting to estimate future spending. Activating cost allocation tags lets you map costs to business units.
Trusted Advisor provides recommendations across 5 categories: cost optimization, performance, security, fault tolerance, and service limits. It is particularly useful for identifying unused resources like idle EC2 instances and unattached EBS volumes.
Compute Optimizer analyzes usage patterns of EC2 instances, Auto Scaling groups, Lambda functions, and EBS volumes to provide right-sizing recommendations. It uses CloudWatch metric data to identify over-provisioned and under-provisioned resources.
Purchasing Options Comparison
| Option | Discount | Commitment | Best For | |--------|----------|------------|----------| | On-Demand | None | None | Unpredictable short-term workloads | | Reserved Instance | Up to 72% | 1yr/3yr | Steady, predictable workloads | | Savings Plans (Compute) | Up to 66% | 1yr/3yr | Region/family flexibility needed | | Savings Plans (EC2 Instance) | Up to 72% | 1yr/3yr | Fixed region+family for max discount | | Spot Instance | Up to 90% | None | Interruptible batch/stateless workloads |
Savings Plans are more flexible than Reserved Instances. Compute Savings Plans apply to EC2, Fargate, and Lambda, and the discount persists even when you change regions or instance families. EC2 Instance Savings Plans lock in a specific region and instance family in exchange for a higher discount.
!EC2 purchasing options compared
S3 Storage Lens analyzes S3 usage patterns across your entire organization. It visualizes per-bucket storage usage, request patterns, and costs, and includes anomaly detection.
Practical Scenarios
Scenario 1: A financial services company requires RPO of 5 minutes and RTO of 30 minutes. Backup-Restore has an RTO of hours, so it is unsuitable. Pilot Light (RDS Cross-Region Read Replica + AMI-based recovery) is the cost-effective choice.
Scenario 2: An e-commerce company requires zero downtime. Multi-Site Active-Active (Route 53 latency-based routing + full stack in two regions) is needed. It has the highest cost but achieves near-zero RTO/RPO.
Scenario 3: A startup needs to reduce monthly AWS costs by 30%. Use Compute Optimizer to identify over-provisioned EC2 instances, apply Compute Savings Plans to steady workloads, and use Spot Instances for batch jobs.
Exam Key Points
"Lowest cost DR" -- Backup and Restore
"Only core DB always on, rest recovered on failure" -- Pilot Light
"Scaled-down full environment always running" -- Warm Standby
"Zero downtime, instant failover" -- Multi-Site Active-Active
"Shorter RTO means" -- higher cost
"On-premises server replication DR to AWS" -- Elastic Disaster Recovery (DRS)
"Identify over-provisioned EC2" -- Compute Optimizer
"Map costs to business units" -- Cost allocation tags + Cost Explorer
"Discount for EC2/Fargate/Lambda, region flexible" -- Compute Savings Plans
"Analyze S3 usage patterns org-wide" -- S3 Storage Lens
"Identify unused resources" -- Trusted Advisor