High availability (HA) is about designing your service to be "almost always" on, while disaster recovery (DR) is about bringing it back quickly if it ever goes down. The AZ-305 exam regularly tests your ability to choose the right layer and the right service for each goal. In this guide, we'll walk through the key services one by one and clarify the selection criteria that trip people up most often.
---
RTO, RPO, and Business Continuity Design
Every disaster recovery design starts with two numbers: RTO and RPO.
RTO (Recovery Time Objective) is the maximum amount of time your service is allowed to be down before it must be back online. If your RTO is four hours, you must restore service within four hours of an outage. A shorter RTO demands faster failover infrastructure — and therefore higher cost.
RPO (Recovery Point Objective) is the maximum tolerable data loss, expressed as a window of time. An RPO of 24 hours means you can afford to lose up to 24 hours worth of data. This value directly determines how frequently you must back up or replicate data. With daily backups, your RPO is at most 24 hours. Switch to hourly backups and RPO shrinks to one hour.
One of the most common mistakes on the exam is confusing retention period with RPO. Retention period answers the question "how far back in time can I restore?" — for example, 12 months of monthly backups means you can restore a snapshot from 10 months ago. RPO answers a completely different question: "how much data could I lose if a failure happened right now?" With daily backups, that answer is still up to 24 hours regardless of how long you keep those backups.
The core principle of business continuity design is matching your recovery approach to your RTO and RPO targets. Loose targets — say, RTO of 12 hours and RPO of 24 hours — can usually be met with a solid backup strategy alone. Tight targets — RTO of minutes and RPO of seconds — require continuous replication and automatic failover.
---
Availability Zones and Availability Sets
Azure offers two foundational ways to protect workloads from single datacenter failures.
An Availability Set places VMs across different physical racks within the same datacenter. VMs are spread across fault domains (separate power and network paths) and update domains (staggered maintenance windows) so that a single hardware failure or a planned update cycle never takes down all your VMs simultaneously. The SLA is 99.95%.
An Availability Zone goes further. Within a single Azure region, each zone is a physically separate datacenter with its own independent power supply, cooling, and network infrastructure. Even if an entire zone suffers a power outage, VMs in the other zones keep running. The SLA is 99.99%.
Availability Zones protect against failures within a region. If the entire region becomes unavailable, Availability Zones are not enough — you need a multi-region architecture.
Exam rule of thumb: rack or host failure within a single datacenter → Availability Set; zone-level failure within a region → Availability Zone; entire region failure → multi-region design.
!Availability Zone versus Availability Set
Region Pairs and Multi-Region Architecture
Azure groups certain regions into geographically close pairs — known as Region Pairs. For example, Korea Central is paired with Korea South.
Region Pairs have three important properties:
Azure platform updates are never deployed to both regions in a pair simultaneously, which distributes the risk of update-related disruptions. If you configure GRS (geo-redundant storage) or cross-region restore (CRR), your data is automatically replicated to the paired region. Azure Key Vault automatically fails over to the paired region during a regional outage and operates in read-only mode. Existing key retrieval, encryption, and decryption continue to work, but creating or modifying keys is not possible during the failover period.
For workloads that must survive a full regional outage, two main multi-region patterns exist:
The first is Active-Active or Warm Standby — running identical infrastructure in the secondary region at all times. This keeps RTO and RPO extremely low, but you pay for the secondary infrastructure continuously.
The second is Cold Standby — keeping only replicated data in the secondary region and provisioning infrastructure on demand when a failure occurs. This is more cost-efficient but results in a somewhat longer RTO. Azure Site Recovery-based VM DR is the classic example of this pattern.
For VMSS-based web back-ends, the recommended pattern is to deploy an identical VMSS in the secondary region and configure global failover using Front Door. Front Door probes backend health every 30 seconds and automatically shifts traffic to the secondary region if the primary becomes unavailable.
---
Azure Site Recovery and Backup Strategy
Azure Backup and Azure Site Recovery sound similar, but they serve very different purposes.
Azure Backup stores point-in-time snapshots of your data in a Recovery Services Vault. You can configure daily, weekly, monthly, and yearly retention tiers, with retention up to 99 years. To back up files and folders from an on-premises Windows Server directly to Azure without a separate backup server, you install the MARS Agent (Microsoft Azure Recovery Services Agent). The MARS Agent sends files, folders, and system state to the Recovery Services Vault independently — it can run alongside the existing Windows Server Backup without conflicts.
Azure Site Recovery (ASR) continuously replicates VM disks to a secondary region, achieving an RPO of a few minutes to around 15 minutes. In normal operation, no VM is running in the secondary region — only the replication data is maintained — which keeps costs low. When a failure occurs, VMs are automatically provisioned in the secondary region, enabling recovery within an RTO of under one hour. ASR supports on-premises to on-premises, on-premises to Azure, and Azure to Azure scenarios. Recovery Plans let you define failover order and handle dependencies automatically. A Test Failover feature lets you validate your DR plan in an isolated network without any impact on production.
The Recovery Services Vault is the central store for both Azure Backup and Azure Site Recovery. Enabling Immutability Lock prevents backup data from being deleted or modified, which is your primary defense against ransomware attacks. With GRS and cross-region restore (CRR), you can restore from the secondary region even during a regional outage. Because a Vault is tied to the region where it was created, multi-region environments require a separate Vault per region. To monitor multiple Vaults from a single pane of glass, use Backup Center.
To apply a standardized backup policy to hundreds of VMs in bulk, link an Azure Policy to your Recovery Services Vault backup policy. New VMs that fall within the policy scope are automatically enrolled, keeping your compliance posture intact without manual intervention.
---
Service Comparison Table
The table below compares the major Azure components for high availability and disaster recovery by protection scope, SLA, and RTO/RPO.
| Component | Protection Scope | SLA | RPO | RTO | Key Characteristics | |-----------|-----------------|-----|-----|-----|--------------------| | Availability Set | Racks and hosts within a datacenter | 99.95% | N/A | N/A | Fault domain distribution, no extra cost | | Availability Zone | Independent datacenters within a region | 99.99% | N/A | N/A | Separate physical power and cooling | | Region Pair + GRS | Cross-region storage replication | 99.99%+ | Minutes to 1 hour | Hours | Storage-tier geo-redundancy | | Azure Backup | Point-in-time data retention | - | Hours to 24 hours | Hours to 12 hours | Long-term retention, PITR | | Azure Site Recovery | Cross-region VM-level replication | - | Minutes to 15 minutes | Under 1 hour | Automatic failover, cost-efficient | | SQL Auto-Failover Group | Azure SQL DB cross-region | 99.99% | Under 5 seconds | Under 1 hour | Single endpoint, automatic switchover | | Always On AG + ILB | VM-based SQL Server | 99.99%+ | 0 seconds (synchronous) | Seconds | IaaS environment, WSFC-based |
---
Selection Criteria That Commonly Confuse Exam Takers
Here is a quick breakdown of the scenarios and decision rules that appear most frequently on the real exam.
Scenario 1: RTO of 12 hours, RPO of 24 hours, secondary infrastructure cannot run continuously, cost efficiency is the priority. Choose Azure Backup with PITR. You do not need to keep VMs running in the secondary region, and daily backups deliver the required RPO of 24 hours. Active Geo-Replication is too expensive for this requirement.
Scenario 2: Automatic failover on regional outage, service must resume without changing the application connection string (Azure SQL Database). Choose Auto-Failover Group. It provides two fixed DNS endpoints — a Primary Listener (read/write) and a Secondary Listener (read-only) — so the connection string never needs to change after a failover. Active Geo-Replication supports manual failover only and requires separate endpoint management.
Scenario 3: SQL Server running on IaaS VMs, automatic failover with no manual intervention. Choose Always On AG + ILB. In synchronous-commit mode, RPO is 0 and automatic failover is supported. Azure SQL DB automatic failover is a PaaS service and does not satisfy the IaaS requirement.
Scenario 4: VMSS-based web service must remain available even during a regional outage. Choose VMSS in the secondary region combined with Front Door. Spreading VMs across Availability Zones alone cannot protect against a full regional failure.
Scenario 5: Agent-based backup of an on-premises file server, no separate backup server allowed. Choose MARS Agent with Recovery Services Vault. The difference from Windows Admin Center is whether an extra agent must be installed. If agent depl