Mastering Monitoring and Logging

Master the DOP-C02 Monitoring domain (15%): CloudWatch collection pipelines, X-Ray distributed tracing, EventBridge-driven automation, and Lambda-based self-healing architectures through real exam scenarios.

The Monitoring and Logging domain accounts for 15% of the DOP-C02 exam. The questions go beyond asking which service collects what data — they evaluate whether you can design a complete pipeline from collection through analysis to automated action. Think of it like an on-call engineer troubleshooting a production incident at 3 AM: the exam asks which combination of tools is right for each situation.

 

Log Collection Pipeline Design

Effective monitoring starts with collecting data correctly. CloudWatch is a powerful platform where a single agent can simultaneously collect logs and custom metrics.

CloudWatch Agent and Custom Metrics

The default CloudWatch metrics for EC2 do not include memory utilization or detailed disk I/O. To collect these, you must install the CloudWatch Agent. The agent works identically on EC2 (Linux and Windows) and on-premises servers.

Consider a gaming company operating thousands of EC2 instances. They were maintaining a homegrown log collection daemon and experiencing frequent configuration gaps when new instances launched. The fix is CloudWatch Agent combined with AWS Systems Manager State Manager. State Manager applies the agent automatically to new instances, while SSM Parameter Store holds the centralized, versioned agent configuration.

| Collection Method | Target | Key Characteristics | |------------------|--------|-------------------| | CloudWatch Agent | EC2, on-premises | Memory/disk custom metrics and log file collection | | CloudWatch Embedded Metric Format (EMF) | Lambda, ECS, EC2 | Embeds metrics in JSON logs, no direct PutMetricData calls | | AWS Distro for OpenTelemetry (ADOT) | Containers, serverless | OpenTelemetry standard, sends to both X-Ray and CloudWatch |

EMF is effective for reducing API costs. Sending multiple metric data points individually via PutMetricData incurs per-call charges. EMF sends structured JSON logs to CloudWatch Logs and AWS automatically extracts the metrics. The cost is lower than direct PutMetricData calls, and the extracted metrics support alarms just like standard CloudWatch metrics.

 

Kinesis Data Firehose with Lambda Transformation

A common exam scenario involves an IoT platform receiving logs in different formats from thousands of devices, needing to store them in S3 and query with Athena. The answer is Kinesis Data Firehose with a Lambda transformation.

Kinesis Data Firehose is a fully managed service requiring no server administration. A Lambda function acts as a data transformation processor, converting various log formats into JSON before writing to S3. Buffer size (1-128 MB) and buffer interval (60-900 seconds) settings provide near-real-time batch delivery.

The distinction from Kinesis Data Streams is clear. Data Streams requires shard management and checkpoint implementation, increasing operational complexity. Firehose delivers automatically to S3, Redshift, or OpenSearch without any consumer code. Firehose is also less expensive.

Failed transformation records are preserved in a separate S3 bucket, ensuring zero data loss. This "error record preservation" behavior is tested when the question includes the condition "no data loss is acceptable."

 

Account-Level Subscription Policies and Cross-Account Log Aggregation

When hundreds of log groups are added dynamically, setting subscription filters individually is operationally impractical. CloudWatch Logs account-level subscription policies, configured once, automatically apply to all current and future log groups in the account. This is the answer pattern when the question combines "minimize operational overhead" with a growing number of log groups.

Multi-account log aggregation architecture:

Each member account: configure a CloudWatch Logs account-level subscription policy Central Audit account: Kinesis Data Firehose waiting to receive logs Member account subscription filters forward directly to the central Firehose (cross-account delivery) S3 Lifecycle Policy: automatic transition to S3 Glacier after 90 days

The Export Task versus subscription filter distinction matters. Export Tasks are batch and manual, with delays up to 12 hours. Subscription filters are streaming and automatic, providing near-real-time delivery. For long-term storage cost optimization, an S3 Lifecycle Policy transitioning to Glacier is essential.

 

ALB Access Log Timing Fields

ALB access logs break request processing time into three fields. The exam asks you to distinguish them precisely.

| Field | Meaning | What High Values Indicate | |-------|---------|--------------------------| | request_processing_time | Time from receiving client request to forwarding to target | ALB processing bottleneck | | target_processing_time | Time for the target (EC2/container) to process the request | Application or database performance issue | | response_processing_time | Time from receiving target response to sending to client | Network or ALB return path latency |

When target_processing_time is high, the problem is in the application code or database queries. When request_processing_time is high, the ALB itself is the bottleneck. These fields are analyzed using CloudWatch Logs Insights or Athena.

 

Distributed Tracing and Hybrid Monitoring

X-Ray Distributed Tracing

In a microservices architecture, when a particular API call is slow, identifying which service is the bottleneck is difficult. X-Ray visualizes the complete path a request takes through multiple services.

Key X-Ray concepts: Trace: the complete execution path of a single request Segment: the unit of work done in each service Subsegment: granular work units such as external API calls and database queries Service Map: visual representation of service dependencies and response times

You can instrument Lambda, ECS, EC2, API Gateway, and Elastic Beanstalk using the X-Ray SDK. Lambda supports X-Ray Active Tracing, which enables basic tracing without code changes.

Amazon Managed Grafana + Prometheus for Hybrid Environments

For unified monitoring of hybrid environments mixing EKS, ECS, and on-premises, the Amazon Managed Grafana + Amazon Managed Service for Prometheus combination appears on the exam. This pair is fully managed with no server administration while being compatible with the Prometheus ecosystem.

The metric collection model differs from CloudWatch. CloudWatch uses a push model; Prometheus uses a pull model, scraping targets for metrics. Grafana supports both data sources, enabling unified dashboards that combine CloudWatch metrics and Prometheus metrics.

 

CloudWatch Logs Analysis Tools

Logs Insights and Metric Filters

CloudWatch Logs Insights provides interactive log analysis using a SQL-like query language with near-real-time search capabilities. This makes it suitable for querying application logs from Elastic Beanstalk, ECS, and other services in real time. Athena is a batch query service that analyzes data already stored in S3.

CloudWatch Logs Metric Filter detects regex patterns in log streams in real time and converts matches into custom CloudWatch metrics. It is configured entirely in the console without any code, keeping operational overhead low.

Security Operations Center scenario: when firewall logs must trigger immediate alerts on CRITICAL severity events without additional security services, the answer is CloudWatch Logs Metric Filter + CloudWatch Alarm + SNS. This three-service pattern is the answer whenever the conditions include "no code required," "minimize operational overhead," and "based on existing CloudWatch Logs."

Metric Filter versus Subscription Filter distinction: Metric Filter: pattern matching produces count metrics that trigger Alarms (for alerting) Subscription Filter: streams log events to Lambda, Kinesis, or Firehose in real time (for data pipelines)

 

Event-Driven Automation and Self-Healing

EventBridge Event Patterns

EventBridge is the hub for building automation chains from events emitted by AWS services. Important event patterns for the exam:

| Service | source Value | Key Events | |---------|-------------|-----------| | EC2 Auto Scaling | aws.autoscaling | EC2_INSTANCE_LAUNCH_UNSUCCESSFUL | | AWS Trusted Advisor | aws.trustedadvisor | Service limit approaching 80% | | AWS Health | aws.health | Instance retirement, service disruption | | AWS Config | aws.config | Configuration compliance change |

Auto Scaling instance launch failure scenario: During a promotional event, Auto Scaling failed to launch instances but the team only found out hours later. The fix is an EventBridge rule that detects EC2_INSTANCE_LAUNCH_UNSUCCESSFUL events and routes them to SNS for immediate notification. No additional infrastructure is needed beyond EventBridge and SNS.

Do not confuse this with Lifecycle Hooks. Lifecycle Hooks run after a successful instance launch while the instance is in the Pending state. They cannot detect a launch failure.

 

CloudWatch Alarms + Lambda for Auto-Remediation

CloudWatch Alarms go beyond simple notifications. They can trigger Lambda functions to execute automated remediation actions.

SSH direct login detection and automatic isolation scenario: A company prohibits SSH and requires Session Manager access only, but SSH login records appear in OS logs on some instances. Implementation:

CloudWatch Agent collects /var/log/secure on Linux instances CloudWatch Logs Metric Filter detects SSH login patterns using regex CloudWatch Alarm triggers Lambda when the count exceeds threshold (1 or more) Lambda replaces the instance's security group with an isolation group or terminates the instance SNS sends immediate notification to the security team

CloudTrail cannot detect SSH logins. CloudTrail records AWS API calls, not OS-level SSH sessions.

AWS Config Rules and Auto-Remediation

AWS Config Rules continuously evaluate resource compliance. Remediation Actions linked to a rule automatically run an SSM Automation runbook when a violatio

Back to blog list