Performance Analysis and Troubleshooting

EC2 instance types, EBS volumes, RDS Performance Insights, CloudTrail, and SSM Automation explained from scratch for beginners.

Troubleshooting performance in AWS is like being a car mechanic. When a car is slow, you need to figure out whether the problem is the engine, the tires, or the fuel. In AWS, when a server is slow, you check whether the problem is CPU, memory, storage, or the network. This post teaches you the diagnostic tools.

 

EC2 Instance Types — Choosing the Right Tool

EC2 instances come in different types, each optimized for a specific kind of work. Think of it like choosing the right kitchen equipment: you would not bake bread in a stockpot or boil soup in an oven.

| Type | Name | Characteristic | Good For | |------|------|----------------|---------| | General Purpose | M series, T series | Balanced CPU and memory | Web servers, small databases | | Compute Optimized | C series | Extra high CPU performance | Batch processing, scientific computing, gaming servers | | Memory Optimized | R series, X series | Very large amounts of RAM | In-memory databases, large caches | | Storage Optimized | I series, D series | Very fast disk read/write | Data warehouses, distributed file systems |

The T Instance Credit System

T series instances (T3, T4g, etc.) work in a unique way. When CPU usage is low, the instance accumulates "credits" over time. When CPU demand spikes, it spends those credits to burst to higher performance. Think of it like a bank account: save up when business is slow, spend when you need it.

When the credit balance hits zero, CPU performance is throttled down to a baseline level (for example, a t3.micro is limited to about 10% CPU). At this point, your application can suddenly become very slow.

Enabling Unlimited mode lets the instance continue running at high CPU even after credits are exhausted, but you are charged for the excess CPU used.

Exam tip: if a question says "an EC2 instance is running slowly" and it is a T-type instance, suspect credit exhaustion first.

Placement Groups — Where to Physically Place Your Servers

Placement groups control the physical location of your instances relative to each other.

Cluster placement group: places all instances close together on the same physical hardware in the same Availability Zone. Network latency between instances becomes extremely low. Use this for machine learning training or HPC workloads where servers need to exchange enormous amounts of data at high speed. The downside is that all instances are in one location, so a failure in that AZ affects them all.

Spread placement group: places each instance on separate physical hardware. If one server's hardware fails, the others are unaffected. Use this for critical instances that must remain isolated from each other for high availability.

Partition placement group: divides instances into groups called partitions, and each partition is on a separate rack of hardware. Use this for large distributed systems like HDFS (Hadoop) or Cassandra, where you want rack-level isolation.

 

EBS Volume Types — Not All Hard Drives Are Equal

EBS (Elastic Block Store) is like a hard drive you attach to your EC2 instance. Different volume types have very different performance and cost characteristics.

| Volume Type | Max IOPS | Max Throughput | Key Feature | Use Case | |------------|---------|---------------|------------|---------| | gp3 | 16,000 | 1,000 MB/s | IOPS configurable independently of size | Most general workloads | | io2 Block Express | 256,000 | 4,000 MB/s | Highest performance, Multi-Attach | Mission-critical databases | | st1 | N/A | 500 MB/s | Sequential throughput optimized, low cost | Big data, log processing | | sc1 | N/A | 250 MB/s | Cheapest option | Archives, infrequently accessed data |

IOPS (Input/Output Operations Per Second) matters for workloads like databases that do many small random reads and writes. Throughput (MB/s) matters for workloads like video processing that read large files sequentially.

The gp3 key point: with the older gp2, IOPS was tied to volume size (more GB = more IOPS). With gp3, you configure IOPS and throughput independently of volume size. This means you do not have to buy more storage just to get more performance.

 

RDS Performance Insights — The Health Report for Your Database

Imagine you run an online store and order processing suddenly slows down. You need to know: is it the network, the servers, or a specific database query? RDS Performance Insights is a visual dashboard that shows you exactly what is happening inside your database engine.

Key features:

Top SQL queries ranked by the load they create. You can see "this one query is responsible for 40% of the total database load."

Wait events show you why queries are waiting: is it because the CPU is overloaded, the disk is slow, or another query is holding a lock on the same data?

DB load graph visualizes how busy the database is compared to the number of vCPUs available. This helps you decide whether to scale up to a larger instance.

How is it different from Enhanced Monitoring? Performance Insights looks inside the database engine itself: queries, execution plans, wait events. Enhanced Monitoring looks at the server running the database: OS-level CPU, memory, and process information. The two tools complement each other.

 

S3 Transfer Acceleration — Fast Uploads from Anywhere in the World

Imagine a team in Seoul needs to upload large video files to an S3 bucket in the US East region. Using the normal internet, the data travels through many intermediary routers and undersea cables, which introduces latency and instability.

S3 Transfer Acceleration solves this by routing the upload to the nearest CloudFront edge location first (Seoul has one). From there, the data travels over AWS's private backbone network, which is much faster and more reliable than the public internet, all the way to the destination S3 bucket in the US.

To use it, change your S3 endpoint to . Combining it with multipart uploads (splitting large files into chunks) makes it even faster.

Important note: Transfer Acceleration is most effective for long-distance uploads across continents. For short distances, it may not improve speed and could even be slower due to the detour through an edge location.

 

EFS Performance Modes — Tuning Your Shared File System

EFS (Elastic File System) is a shared folder that multiple EC2 instances can access simultaneously. Think of it like a shared network drive in an office, except it can scale to petabytes and thousands of concurrent connections.

There are two performance modes. General Purpose mode has lower latency and is suitable for web servers, content management systems, and home directories where fast response time matters more than raw throughput. Max I/O mode has higher throughput but slightly higher latency. It is designed for big data analytics and media processing where thousands of instances access the file system simultaneously.

For throughput, there are three modes. Bursting throughput scales automatically based on file system size. Provisioned throughput lets you specify exactly how much throughput you need regardless of size. Elastic throughput automatically adjusts to your actual workload without any configuration.

 

CloudTrail — The Audit Log for Your Entire AWS Account

CloudTrail records every API call made in your AWS account. Think of it as the access badge system for a secure office building: every time someone opens a door, it is recorded with a timestamp and employee name.

When a security incident occurs, CloudTrail lets you answer questions like: "Who made this S3 bucket public?", "When was this IAM user created?", "Which IP address ran this command?"

Event types:

Management events record operations on your infrastructure: launching EC2 instances, creating IAM roles, creating S3 buckets. These are collected by default and are free.

Data events record operations on the data inside your resources: reading and writing S3 objects, invoking Lambda functions. These are not collected by default and have an additional cost.

Insights events automatically detect unusual API activity patterns. For example, if there is suddenly a huge spike in IAM role creation API calls, CloudTrail Insights flags it as a potential security threat.

CloudTrail Lake lets you run SQL queries directly against your CloudTrail event history. This is useful for compliance audits and forensic investigation.

 

Systems Manager Automation — Stop Doing Repetitive Tasks Manually

Manually SSH-ing into servers to restart them every time something goes wrong is inefficient and error-prone. AWS Systems Manager Automation lets you define runbooks that perform these tasks automatically.

A runbook is a document that defines a sequence of steps, like a recipe. For example, an "EC2 Restart Runbook" might have these steps: step 1 send a notification, step 2 stop the instance, step 3 start the instance, step 4 verify the instance is healthy. All of this happens without human intervention.

The most common exam pattern: a CloudWatch alarm fires because a metric exceeds a threshold. EventBridge detects the alarm state change and triggers an SSM Automation runbook. The runbook restarts the server. The whole thing happens automatically while the on-call engineer sleeps.

AWS provides pre-built runbooks for common scenarios: AWS-RestartEC2Instance, AWS-StopEC2Instance, and others are available out of the box.

 

Exam Key Points

"T instance suddenly slow" -- Check CPU credit balance; enable Unlimited mode or change instance type

"Minimize network latency between servers" -- Cluster placement group

"Isolate critical instances from hardware failures" -- Spread placement group

Back to blog list