Compute Services Part 2

A simple guide to Auto Scaling, ELB types, and EC2 instance families.

What Happens When Traffic Suddenly Spikes?

Imagine you launch a new service and thousands of people try to use it at the same time. With only one server, it would likely crash from the overload.

On the other hand, if you keep ten servers running at all times just in case, you are wasting money every night when almost nobody is online.

AWS Auto Scaling and load balancers solve both problems automatically.

---

 

Think of Highway Traffic Management

Picture a smart highway that automatically opens extra lanes when traffic is heavy and closes them when roads are empty.

Auto Scaling = the smart lane system that adds servers when users arrive and removes them when they leave. Elastic Load Balancing = the traffic officer directing cars evenly across all available lanes.

Used together, these two services ensure your application handles large crowds without crashing, and costs almost nothing during quiet hours.

---

 

Auto Scaling — Automatically Add and Remove Servers

Amazon EC2 Auto Scaling adjusts the number of EC2 instances automatically based on demand.

How does it work? You can configure it so that when CPU usage exceeds 70%, a new server is added. When CPU drops below 20%, a server is removed.

You set three numbers.

The Minimum is the smallest number of servers that must always be running. Even at 3 AM with no users, you never drop below this number.

The Maximum is the highest number of servers you allow. This acts as a cost ceiling so your bill cannot grow beyond a limit you set.

The Desired is the number of servers you want to maintain during normal conditions. Think of it as your everyday baseline.

Auto Scaling has two key benefits. High Availability: if a server fails, Auto Scaling automatically replaces it with a new one. Cost Optimization: you only pay for servers you actually need at any given moment.

---

 

Elastic Load Balancing — Distribute Traffic Evenly

A load balancer takes all incoming user requests and spreads them evenly across your servers.

Why is this needed? Even if you have three servers, without a load balancer all traffic might pile onto just one of them. A load balancer monitors the health of each server and only sends traffic to healthy ones. If a server goes down, the load balancer automatically routes traffic around it.

AWS offers three types of load balancers.

ALB (Application Load Balancer) is the most commonly used type for web applications. It handles HTTP and HTTPS traffic and can even route requests to different servers based on the URL path — for example, sending requests for /images to one group of servers and /api to another.

NLB (Network Load Balancer) is designed for situations requiring extreme speed. It handles TCP and UDP traffic and can process millions of requests per second. It is ideal for online games and real-time streaming services.

GLB (Gateway Load Balancer) is used when you need all traffic to pass through a third-party security appliance — such as a firewall or intrusion detection system — before reaching your servers.

---

 

EC2 Instance Type Families

Choosing the right type of server for your workload is just as important as having the right number of servers.

| Family | Characteristic | Best Used For | |--------|---------------|--------------| | General Purpose | Balanced CPU and memory | Web servers, development environments, mid-size databases | | Compute Optimized | Extra-powerful CPU | Scientific simulations, game servers, video encoding | | Memory Optimized | Extra-large RAM | Large in-memory databases, real-time big data analytics | | Storage Optimized | Very fast disk read/write speeds | Data warehouses, distributed file systems | | Accelerated Computing | GPU (graphics processing unit) | Machine learning training, graphics rendering |

---

 

Dedicated Hosts vs Dedicated Instances

Most AWS customers share the underlying physical hardware with other customers. But sometimes legal, security, or licensing requirements demand exclusive use of hardware.

Dedicated Instances guarantee that no other customer's instances run on the same physical server as yours. However, other instances within your own AWS account can still share the same physical machine.

Dedicated Hosts give you exclusive use of an entire physical server. You can see and control exactly how many sockets and cores the server has. This is primarily used when you need to bring existing software licenses to AWS (BYOL — Bring Your Own License).

---

 

Exam Key Points

"Automatically adjust server count based on traffic" → Amazon EC2 Auto Scaling "Distribute traffic evenly across multiple servers" → Elastic Load Balancing (ELB) "Load balancer for HTTP/HTTPS web traffic" → ALB (Application Load Balancer) "Ultra-low-latency TCP/UDP load balancer" → NLB (Network Load Balancer) "Best for CPU-intensive workloads" → Compute Optimized instances "Best for workloads requiring large amounts of memory" → Memory Optimized instances "Use existing software licenses on AWS (BYOL)" → Dedicated Hosts Auto Scaling plus ELB together = the fundamental AWS pattern for high availability and cost optimization

Back to blog list