High-Performance Compute

Choose EC2 instances, configure Auto Scaling policies, and handle batch/HPC workloads.

In the SAA-C03 exam, compute performance falls under the Design high-performing architectures domain, which makes up roughly 26% of the exam. Questions repeatedly ask: "Which EC2 instance type fits this workload?" and "How do you automatically scale capacity up and down?"

 

EC2 Instance Families — Choosing the Right Vehicle for the Job

Think of EC2 instance families like vehicles. You would not drive a sports car when moving furniture, and you would not use a cargo truck for a Sunday drive. Each EC2 family is engineered for a specific type of workload.

M-series — The Family Sedan (General Purpose)

Balanced CPU-to-memory ratio. Use this for web servers, application servers, and small databases when there is no extreme requirement in any single dimension. If you are not sure what to pick, start with M.

C-series — The Sports Car (Compute Optimized)

Built for CPU-heavy tasks. Scientific simulations, video encoding, gaming servers, and high-traffic web front-ends all benefit from the extra CPU horsepower C-series delivers at the same cost.

R/X-series — The Cargo Van (Memory Optimized)

A cargo van is designed to carry a lot of load. R and X series give you an enormous amount of RAM relative to CPU. Use these for in-memory databases like SAP HANA, real-time big data analytics, and large caching layers.

P/G-series — The Specialized Truck (Accelerated Computing)

When your workload demands GPUs — machine learning model training, deep learning inference, 3D rendering — P and G series are the only choice. They are over-spec for general tasks, but irreplaceable when GPU power is needed.

I/D-series — The High-Speed Warehouse (Storage Optimized)

When disk read/write throughput is the bottleneck, I and D series provide extremely fast local NVMe storage. Ideal for data warehouses, large-scale log processing, and NoSQL database servers.

 

Auto Scaling — Hiring Temp Workers During the Holiday Rush

Imagine a retail store during the holiday shopping season. The week before a major holiday, customer traffic spikes dramatically. The store manager hires extra temporary workers in advance. Once the holiday is over, those workers are let go. But no matter how slow business gets, there is always at least one cashier on duty. Auto Scaling does exactly this, automatically.

Three numbers define an Auto Scaling Group:

Minimum: The floor. Instances never drop below this. "We always need at least one cashier." Desired: The target for normal operations. "Three staff members is the right size for a typical day." Maximum: The ceiling. Never exceed this even during peak load. "We can fit at most ten people in the store."

 

Auto Scaling Policies — What Triggers Scaling?

Target Tracking

The simplest and most commonly used policy. You set a target metric — for example, "keep CPU utilization at 50%" — and AWS automatically adds or removes instances to stay near that target. Think of it like a thermostat: you set the temperature, and the system handles the rest.

Step Scaling

More granular control. You define scaling steps: add 2 instances when CPU is between 60–70%, add 4 when CPU is between 70–80%, and so on. Use Step Scaling when you need finer control than Target Tracking provides.

Scheduled Scaling

Scale on a timetable. "Every weekday at 8 AM, bring capacity up to 10 instances. At midnight, scale back down to 2." Use this when traffic patterns are predictable and consistent.

Predictive Scaling

AWS analyzes your historical traffic patterns using machine learning and proactively adds capacity before demand arrives. Instead of reacting to a traffic spike, Predictive Scaling prepares for it in advance.

 

EC2 Placement Groups — Where Exactly Do Your Servers Live?

Where AWS physically places your instances within its data centers has a significant impact on both performance and availability.

Cluster Placement Group — All Desks in One Office

If everyone on a team sits in the same room, they can communicate at maximum speed with minimum delay. A Cluster Placement Group places all instances on hardware that is physically close together within a single AZ. This minimizes network latency and maximizes throughput between instances. Use it for HPC (high-performance computing) workloads and tightly coupled applications that require fast node-to-node communication. Trade-off: if the underlying hardware rack fails, all instances are affected.

Spread Placement Group — One Desk Per Floor

If each employee sits on a different floor, a fire on one floor does not affect the others. Spread places each instance on distinct physical hardware, maximizing fault isolation. Use this for a small number of critical instances that absolutely cannot fail at the same time. Limit: a maximum of 7 instances per AZ.

Partition Placement Group — Separate Rooms for Each Team

Divide the company into teams (partitions), each with their own room. Instances within a partition may share hardware, but partitions never share hardware with each other. Use this for large distributed systems like Hadoop, Kafka, and Cassandra that are partition-aware and can handle a subset of nodes failing independently.

!EC2 placement groups: Cluster, Spread, Partition

AWS Batch and EMR — Large-Scale Batch Processing

AWS Batch

You define the work — for example, "process these 10,000 simulation jobs" — and AWS Batch automatically provisions the right EC2 instances, distributes the jobs across them, monitors progress, retries failures, and terminates instances when everything is done. You only pay for the compute time you actually use.

Amazon EMR

EMR lets you run open-source big data frameworks like Hadoop, Spark, Hive, and Presto on a managed EC2 cluster. Use it for large-scale log analysis, data transformation (ETL), and machine learning preprocessing.

 

Exam Key Points

"CPU-intensive, scientific computing, gaming servers" -- C-series (Compute Optimized)

"In-memory database, large RAM requirement" -- R/X-series (Memory Optimized)

"GPU, machine learning training, deep learning" -- P/G-series (Accelerated Computing)

"General-purpose, no extreme requirements" -- M-series (General Purpose)

"Maintain a specific CPU percentage target" -- Target Tracking Auto Scaling

"Different response per metric range" -- Step Scaling

"Scale at a specific time of day" -- Scheduled Scaling

"Proactively scale based on historical patterns" -- Predictive Scaling

"HPC, low latency, tight node-to-node communication" -- Cluster Placement Group

"Small number of critical instances, max fault isolation" -- Spread Placement Group

"Hadoop/Kafka/Cassandra, partition-aware distributed systems" -- Partition Placement Group

"Automated scheduling of large batch jobs" -- AWS Batch

"Hadoop/Spark big data cluster" -- Amazon EMR

Back to blog list