In the AWS DOP-C02 exam, the domains covering "High Availability and Resilience Design," "Scalability Solutions," and "Automated Recovery and Disaster Recovery" together account for roughly 35% of the total questions. Rather than pure memorization, the exam tests architectural judgment — which means you need to understand both the underlying mechanics of each service and the exact scenarios where each applies.
High Availability Design Patterns
ALB Deep Health Checks — Automatic SPOF Isolation
The default ALB health check only verifies whether a TCP port is open. This means the load balancer will continue routing traffic to an instance that is running but cannot actually serve requests because its database connection or an external dependency has failed. Users experience 500 or 503 errors as a result.
The solution is to configure ALB to use an HTTP health check against a dedicated endpoint (typically ) that proactively tests internal dependencies.
When ALB receives a 503 from the health check, it automatically removes the instance from the target group. This requires only an application code change — no infrastructure modification — making it the correct answer to exam questions that specify "minimize changes to existing infrastructure."
Route 53 Failover operates at the DNS level and cannot control individual instance traffic with the same granularity. CloudWatch Alarm combined with Auto Scaling works by replacing unhealthy instances, which introduces latency compared to immediately blocking traffic at the load balancer.
Aurora Global Database — Global Reads with Regional Data Sovereignty
Aurora Global Database consists of one primary region (read-write) and up to five secondary regions (read-only). It uses a dedicated storage-layer replication infrastructure, achieving sub-second replication latency from primary to secondary regions.
| Attribute | Aurora Global Database | DynamoDB Global Tables | |-----------|------------------------|------------------------| | Data model | Relational (MySQL/PostgreSQL) | NoSQL (key-value/document) | | Write regions | Single primary region only | All regions simultaneously | | Replication mechanism | Dedicated physical replication layer | DynamoDB Streams-based | | Disaster recovery | Secondary region promotion under 1 minute | Automatic active-active failover |
The Write Forwarding feature allows writes issued against a secondary region to be automatically forwarded to the primary region for processing. While convenient, the data is actually stored in the primary region — not the secondary. In environments with regional data sovereignty requirements, Write Forwarding violates compliance and must not be used.
DynamoDB Global Tables — Multi-Region Active-Active
When all regions in a global service must handle both reads and writes simultaneously, DynamoDB Global Tables is the correct choice.
Key characteristics: All region replicas accept both reads and writes Asynchronous replication based on DynamoDB Streams, typically propagating within one second Concurrent write conflicts resolved automatically using Last Writer Wins (LWW) No custom replication pipeline required
The distinction from Aurora Global Database is a frequent exam topic. The keyword "writes from all regions" points to DynamoDB Global Tables. The keyword "maintain relational database structure" points to Aurora Global Database.
Route 53 Latency Routing + Health Check
To achieve both latency optimization and automatic failover simultaneously, attach Health Checks to Latency routing policy records — do not use the Failover record type.
Latency routing: directs each user to the region with the fastest response time Health check association: automatically excludes failed regions from DNS responses
The Failover record type is a binary Primary/Secondary construct. During normal operation it cannot distribute traffic across multiple regions. Geolocation routing serves users based on their geographic location, which suits compliance and content localization use cases. For performance optimization, Latency routing is the right choice.
ARC Zonal Shift — Immediate AZ Isolation
When infrastructure fails in a specific AZ, traffic may continue flowing into that AZ while Auto Scaling is still in the process of replacing unhealthy instances. AWS Application Recovery Controller's Zonal Shift moves all ALB or NLB traffic from the affected AZ to healthy AZs with a single API call.
Immediate traffic shift with no DNS TTL delay Integrates with ALB and NLB at the target group level Can be triggered automatically based on CloudWatch alarms Can remain active for up to 72 hours
Route 53 failover requires waiting for the DNS TTL to expire before propagation completes. Zonal Shift is the correct answer whenever the scenario demands immediate per-AZ traffic isolation.
Scalability Solutions — Core Patterns
ASG Lifecycle Hooks + Warm Pool — Solving Initialization Delays
When instances require several minutes after boot to complete initialization (downloading a machine learning model, installing a security agent, starting an application), registering them with the load balancer before initialization finishes causes user-facing request failures.
How Lifecycle Hooks work:
The default wait time is one hour, extendable to 48 hours. Sending an ABANDON signal terminates the instance if initialization fails.
Warm Pool maintains pre-initialized instances in stopped or running state. When a scale-out event occurs, Warm Pool instances transition to InService immediately, eliminating cold-start delays. Instances in stopped state incur only EBS storage costs — no EC2 instance charges — making this approach cost-efficient.
Lifecycle Hook combined with Warm Pool: Lifecycle Hook: blocks load balancer registration until initialization completes Warm Pool: eliminates cold-start delay by keeping pre-initialized instances ready
Auto Scaling Policy Combination Strategy
| Policy | Behavior | Best Use Case | |--------|----------|---------------| | Target Tracking | Maintains a metric target value, AWS manages alarms automatically | General workloads | | Step Scaling | Immediate scaling steps at CloudWatch alarm thresholds | Unpredictable traffic spikes | | Scheduled Scaling | Pre-scales at defined times | Predictable traffic patterns | | Predictive Scaling | ML-based, requires 7–14 days of historical data | Recurring cyclical patterns |
Exam scenarios asking to handle "unpredictable traffic spikes" while "reducing costs during nights and weekends" — without adding Lambda or EventBridge — call for the Target Tracking + Step Scaling combination. Predictive Scaling struggles with one-off event spikes. Scheduled Scaling only helps when the pattern is consistent.
Kinesis Data Streams + Lambda + DynamoDB Serverless Pipeline
This is the standard serverless pipeline design for processing high-frequency events from thousands of IoT sensors or similar sources in real time.
Kinesis Data Streams scales via shards. Each shard handles 1 MB/s writes and 2 MB/s reads. Lambda event source mapping automatically triggers Lambda execution, with concurrency proportional to the shard count.
Kinesis Data Firehose vs. Kinesis Data Streams: Kinesis Data Streams: real-time processing, requires a custom consumer application, millisecond latency Kinesis Data Firehose: fully managed, delivers directly to S3/Redshift/OpenSearch, may buffer for minutes
If key-value storage and real-time dashboard lookups are needed, use Kinesis Data Streams + Lambda + DynamoDB. If long-term archival and batch analytics are the goal, use Kinesis Data Firehose + S3.
ECS + Fargate — Serverless Containers
To eliminate the operational burden of OS patching, security updates, and capacity management in an EC2-based container environment, use AWS Fargate.
Fargate: AWS fully manages the container host ECS + Fargate: appropriate for straightforward container workloads and small teams EKS + Fargate: appropriate when Kubernetes features are required
ECS service integrations with ALB support dynamic port mapping, allowing multiple container tasks to be registered in an ALB target group simultaneously.
When ECS Fargate deployed in a private subnet fails to pull images from ECR, and IAM permissions are confirmed correct, the root cause is almost always missing VPC endpoints. Required endpoints:
— ECR API calls — Docker registry protocol (gateway) — ECR image layers are stored in S3 — CloudWatch Logs (recommended)
Correct IAM permissions combined with image-pull failure is a strong signal of missing VPC endpoints.
Disaster Recovery Strategy Comparison
The Four DR Strategies
| Strategy | RTO | RPO | Cost | Description | |----------|-----|-----|------|-------------| | Backup & Restore | Hours | Hours | Lowest | Full restoration from backup | | Pilot Light | 30–60 min | Minutes–seconds | Low | Core services running at minimal scale | | Warm Standby | 5–15 min | Seconds–minutes | Medium | Reduced-scale full stack always running | | Multi-Site Active-Active | Seconds or less | Near-zero | Highest | All regions serving live traffic |
Exam questions provide RTO and RPO numbers as the primary selection criteria: RTO 15 min, RPO 5 min: Warm Standby is optimal (Pilot Light's scale-up time makes 15-minute RTO uncertain) RPO under 1 second, RTO under 1 minute: requires Multi-Site Active-Active or Aurora Global Database Reasonable cost with RTO in tens of minutes: Warm Standby
!The 4 disaster recovery strategies
AWS Elastic Disaster Recovery (AWS DRS)
AWS DRS is a purpose-built DR service for EC2-based workloads. Its approach differs fundamentally from periodic AMI snapshot copying.
How it works: Install a lightweight replication agent on source servers Continuously replicate at the block level to a staging area in the DR region Staging area uses low-cost EBS storage to minimize replication costs On actual failover, launch full EC2 instances from the replicated da