CloudWatch is AWS's monitoring and observability service. Think of it as a hospital monitoring system for your servers. Just as nurses watch a patient's heart rate, blood pressure, and oxygen levels on a monitor and an alarm goes off when something is wrong, CloudWatch watches your servers' CPU, network, and disk activity and alerts you when something is off.
If you have never used AWS before, here is the simplest way to understand CloudWatch: it is the answer to "how do I know if my AWS resources are healthy right now?"
CloudWatch Metrics — The Numbers That Describe Your System
A metric is a measurement at a specific point in time. For example: "CPU usage is 73% right now" or "the server received 500 MB of data in the last minute." CloudWatch stores these measurements over time so you can see trends and spot problems.
Metrics come in two types.
Standard Metrics vs Custom Metrics
| Type | Standard Metrics | Custom Metrics | |------|-----------------|----------------| | Who collects them? | AWS collects automatically | You send them yourself | | Cost | Free | Charged per metric | | Examples | EC2 CPU usage, network in/out, EBS read/write | Memory usage, disk free space, app response time | | Collection interval | Default 5 minutes (detailed: 1 minute) | As frequent as every 1 second |
Here is a critically important exam point that trips up many people: EC2 memory (RAM) usage and disk free space are NOT included in standard metrics.
Why? Because that information lives inside the operating system of your server. AWS manages the physical hardware outside your server but cannot look inside the OS itself. To get memory and disk data, you must install the CloudWatch Agent software on your EC2 instance. The Agent reads the OS internals and sends the data to CloudWatch as custom metrics.
Think of it like this: AWS can see the outside of your apartment building (is the electricity on? is the building structurally sound?), but it cannot walk into your apartment and check what is inside your refrigerator. The CloudWatch Agent is like giving AWS a key to your apartment.
Detailed Monitoring
By default, CloudWatch collects metrics every 5 minutes. Enabling Detailed Monitoring changes this to every 1 minute. There is an additional cost, but it means you detect problems sooner. This is especially useful when combined with Auto Scaling: instead of waiting 5 minutes to notice a traffic spike and add servers, you can react in 1 minute.
CloudWatch Alarms — Automatic Action When Something Goes Wrong
An alarm is a rule that says "when this metric crosses this threshold, do this action." Think of a smoke detector: when smoke reaches a certain level, the alarm sounds.
Three Types of Alarms
Static threshold alarms compare a metric to a fixed number. For example: "send me an email when CPU usage exceeds 80%." Simple and predictable.
Anomaly Detection alarms use machine learning to learn the normal pattern of a metric over time, then alert when the metric deviates from that pattern. For example, if your server CPU always spikes at 9 AM on weekdays, anomaly detection knows this is normal. But if CPU suddenly spikes at 3 AM, it flags that as unusual. This is powerful when you do not know what a "normal" threshold should be.
Composite Alarms combine multiple alarms using AND or OR logic. For example: "only fire an alarm if CPU is high AND memory is also high." This reduces false alarms (also called alarm noise). Imagine you only want to wake up the on-call engineer if both the CPU and the memory are struggling at the same time, not just one of them.
| Alarm Type | How It Works | Best For | |-----------|-------------|---------| | Static threshold | Compare to a fixed number | When you know what "too high" means | | Anomaly Detection | Learn normal patterns, detect deviations | When patterns are complex or unknown | | Composite Alarm | Combine multiple alarms with AND/OR | Reducing false alarms |
What Actions Can an Alarm Take?
When an alarm fires, it can automatically do one of three things.
Send an SNS notification: this can trigger an email to your team, an SMS message, a Slack message through a Lambda function, or any other notification you set up.
Perform an EC2 action: stop the instance, terminate it, reboot it, or recover it to new hardware. The recovery action is particularly important for the exam. When a server has a hardware failure, the StatusCheckFailed_System metric goes to 1. If you have configured a Recovery alarm action on this metric, AWS automatically moves your instance to a fresh piece of hardware while preserving the IP address, instance ID, and data.
Trigger Auto Scaling: add more servers when traffic is high, or remove servers when traffic is low.
CloudWatch Dashboards — See Everything at Once
A dashboard is a customizable page that displays multiple metrics as graphs and numbers in one place. Think of it as the control room of a power plant, where engineers can see the status of the entire facility on a wall of monitors.
Dashboards are useful because:
You can combine metrics from multiple AWS accounts into a single view. For example, see your development, staging, and production environments side by side.
You can show metrics from multiple AWS regions on the same dashboard.
The auto-refresh interval can be set to 10 seconds, 1 minute, or 5 minutes so your view stays current.
CloudWatch Logs — Storing and Searching Log Data
A log is a text record of what happened in a system. Your web server logs every request: who visited, what page they requested, when they did it, and what response was sent back. When something goes wrong, you search the logs to find the cause.
The Structure of CloudWatch Logs
A Log Group is a container for related logs. For example, all logs from a Lambda function called my-order-processor would go into a log group named /aws/lambda/my-order-processor. Think of it as a folder.
A Log Stream is a sequence of log events within a log group. Each EC2 instance or Lambda invocation creates its own log stream within the group. Think of it as a file inside the folder.
Retention period controls how long logs are kept. The default is forever, but you can set it from 1 day to 10 years to control storage costs.
Metric Filters — Turning Log Text Into Numbers
A metric filter scans your logs for a specific text pattern and counts how many times it appears, turning that count into a CloudWatch metric.
Here is a concrete example: your application logs errors with the word "ERROR" in the message. You create a metric filter that counts every occurrence of "ERROR" in the log group. Then you create a CloudWatch alarm on that metric: if more than 10 errors appear within 5 minutes, send an alert to the team. This gives you automated error rate monitoring without any code changes.
CloudWatch Logs Insights — Query Your Logs Like a Database
Logs Insights lets you search and analyze log data using a query language similar to SQL. Instead of scrolling through thousands of log lines, you can run queries like "show me all requests that took longer than 3 seconds in the past hour, grouped by endpoint." Results can be visualized as charts.
EventBridge — React to Events Automatically
EventBridge is a service that watches for events happening across your AWS environment and automatically takes action based on rules you define. The pattern is always: "when event X happens, do action Y."
Here are real examples of what EventBridge can do:
When an EC2 instance is terminated, immediately send an SNS notification to the operations team.
Every day at midnight, trigger a Lambda function that backs up the database.
When an IAM user logs in from an unusual geographic location, trigger a security review workflow.
EventBridge targets (what it can trigger) include Lambda functions, SNS topics, SQS queues, Step Functions workflows, and many more.
Note: EventBridge was previously called CloudWatch Events. You may see both names in the exam. They are the same service.
Exam Key Points
"Monitor EC2 memory and disk usage" -- Install CloudWatch Agent, collect as custom metrics
"Change collection from 5-minute to 1-minute intervals" -- Enable Detailed Monitoring
"Alert when behavior deviates from normal patterns" -- Anomaly Detection alarm
"Only fire alarm when CPU is high AND memory is high" -- Composite Alarm