Auto Scaling and high availability are about making sure your service never goes down, even when individual components fail.
Think of running a restaurant. During the lunch rush, you call in extra staff (Auto Scaling out). When it slows down in the evening, you send staff home early (Auto Scaling in). You spread customers across all available tables (Load Balancing). If one kitchen catches fire, the other kitchen keeps cooking (Multi-AZ failover). This is the essence of AWS high availability design.
Auto Scaling Group (ASG) — Automatically Adjust Your Server Count
An ASG automatically adds or removes EC2 instances based on demand.
Launch Template — The Recipe for Creating Instances
A Launch Template is like a cooking recipe: it defines what ingredients to use (AMI, instance type) and how to prepare them (security groups, key pair, network settings). It also includes the User Data script that runs automatically when an instance starts, letting you install software or configure settings at boot time.
Launch Configuration is an older method that still works but is no longer recommended. Launch Template has more features: versioning, support for multiple instance types, and Spot configuration.
Capacity Settings
| Setting | Meaning | Example | |---------|---------|---------| | Minimum | Never go below this number | Always keep at least 2 | | Desired | The target number right now | Currently aiming for 5 | | Maximum | Never go above this number | Never exceed 20 regardless of load |
Setting Minimum to 0 allows scaling down to zero (no cost when no traffic), but this is risky for production systems. Most production environments keep a minimum of 2 instances to maintain availability.
Scaling Policies — When to Add and Remove Servers
Target Tracking
The simplest and most recommended scaling policy. You set a target metric value and the ASG automatically adjusts the number of instances to maintain it. For example: "keep average CPU usage at 50%." The ASG adds instances when CPU climbs above 50% and removes them when it drops well below 50%. Think of a thermostat: you set the temperature and the system handles heating and cooling automatically.
Step Scaling
Different alarm thresholds trigger different scale actions. For example: CPU 60–70%: add 1 instance CPU 70–80%: add 2 instances CPU above 80%: add 3 instances
This gives you finer control over the scaling response. When load is high but not extreme, you add a little. When load is very high, you add more aggressively.
Scheduled Scaling
Time-based scaling for predictable traffic patterns. For example: "at 8:30 AM every weekday, set capacity to 10; at 6:00 PM, reduce to 3." This is perfect for internal business applications where usage is high during work hours and almost zero at night.
Predictive Scaling
Uses machine learning to analyze historical traffic patterns and proactively scale before demand arrives. If your application consistently sees a traffic spike every Monday morning, Predictive Scaling starts adding instances Sunday night so they are ready when the rush begins.
| Policy | Mechanism | Best For | |--------|-----------|---------| | Target Tracking | Set a target value, ASG maintains it | Most common workloads | | Step Scaling | Different actions at different thresholds | Fine-grained control | | Scheduled Scaling | Time-based schedule | Predictable traffic patterns | | Predictive Scaling | ML forecasts future demand | Recurring spikes, proactive scaling |
!4 auto scaling policy types
ELB — Distributing Traffic Across Multiple Servers
ELB (Elastic Load Balancer) distributes incoming requests across multiple EC2 instances. Think of a bank: instead of all customers queuing for one teller, they are distributed across multiple teller windows so no one waits too long and no teller is overwhelmed.
Four Types of ELB
ALB (Application Load Balancer) operates at Layer 7 (HTTP/HTTPS). It can route requests to different server groups based on URL path, hostname, HTTP headers, or query parameters. For example: requests to go to your API servers, and requests to go to your content servers. It also supports WebSocket connections. Use ALB for web applications and microservices.
NLB (Network Load Balancer) operates at Layer 4 (TCP/UDP). It handles millions of requests per second with extremely low latency (microseconds). It can have static IP addresses, which is important when clients or firewalls need to whitelist specific IPs. Use NLB for gaming servers, IoT devices, financial trading systems, and any application where every millisecond matters.
CLB (Classic Load Balancer) is the original AWS load balancer. It still works but has fewer features. Do not use it for new systems.
GWLB (Gateway Load Balancer) operates at Layer 3 and is used for inserting network appliances (firewalls, intrusion detection/prevention systems) transparently into your network traffic path. All traffic passes through the security appliance before reaching your application.
| ELB Type | Layer | Key Feature | Use Case | |---------|-------|------------|---------| | ALB | Layer 7 (HTTP) | Path/host-based routing | Web apps, microservices | | NLB | Layer 4 (TCP/UDP) | Ultra-low latency, static IP | Gaming, IoT, financial | | CLB | Layer 4/7 | Legacy | Existing systems only | | GWLB | Layer 3 | Transparent appliance insertion | Firewalls, IDS/IPS |
Health Checks — Knowing When a Server is Broken
A health check periodically tests whether each instance is working. If an instance fails the health check, the load balancer stops sending it traffic and the ASG replaces it with a new instance.
EC2 Health Check vs ELB Health Check
| Type | What It Checks | What It Detects | |------|----------------|-----------------| | EC2 health check | Physical hardware and hypervisor status | Hardware failure: the physical machine is dead | | ELB health check | Application response (HTTP status code) | Application failure: the app is returning errors |
This is one of the most important exam points in this domain: if you configure an ASG with only the default EC2 health check (not ELB health check), here is what can happen. Your application crashes and starts returning HTTP 500 errors. Users cannot use the application. But the EC2 instance itself (the physical machine) is still running normally. The EC2 health check says "healthy." The ASG never replaces the broken instance.
To fix this, enable ELB health checks on the ASG. Now when the application crashes and returns errors, ELB's health check fails, it tells the ASG the instance is unhealthy, and the ASG terminates and replaces it automatically.
Multi-AZ Design — Surviving When One Location Fails
An Availability Zone (AZ) is essentially an independent data center. If one AZ goes down due to a power outage or natural disaster, the other AZs are not affected.
The core principle of Multi-AZ design: spread your resources across at least two AZs.
Configure your ASG to span multiple AZs. If one AZ has problems, instances in the other AZs continue handling traffic.
ALB and NLB automatically distribute traffic across instances in multiple AZs. You do not need to configure anything extra.
RDS Multi-AZ synchronously replicates your database to a standby instance in a different AZ. If the primary database fails, AWS automatically promotes the standby to primary, usually within 60-120 seconds. The endpoint (the address your application connects to) stays the same, so no application code changes are needed.
ElastiCache with multiple AZ replicas both improves read performance (read-replicas handle read requests) and provides failover capability (if the primary fails, a replica is promoted).
Mixed Instances Policy — Cutting Costs with Spot Instances
On-Demand instances are always available at a fixed price but are expensive. Spot instances use AWS's spare capacity at up to 90% discount, but AWS can reclaim them with 2 minutes' notice.
The Mixed Instances Policy lets a single ASG combine both. You use On-Demand for your stable baseline capacity and Spot for additional capacity when you need to scale out.
Example: always keep 4 On-Demand instances running. When traffic increases and you need more, the ASG adds Spot instances. When traffic decreases, the Spot instances are terminated first.
Capacity Rebalancing: when AWS is about to reclaim a Spot instance, it proactively starts a replacement Spot instance before terminating the old one, preventing sudden capacity drops.
Specifying multiple instance types: by listing several similar instance types (m5.large, m5a.large, m4.large), you give the ASG more options in the Spot market, which means it is more likely to find available capacity.
Capacity Reservations
Pre-reserve instance capacity in a specific AZ. Use this when you have compliance requirements or DR plans that guarantee "we must always have 10 c5.xlarge instances available in us-east-1a regardless of Spot market conditions."
Exam Key Points
"Simplest way to keep CPU at 50%" -- Target Tracking
"Add different numbers of instances at different CPU levels" -- Step Scaling