Improving Operational Excellence

CloudWatch composite alarms, Systems Manager Automation, deployment strategy comparison, EventBridge event-driven automation, and AWS Config auto-remediation from a SAP-C02 perspective.

In SAP-C02 Domain 3 (Continuous Improvement for Existing Solutions), operational excellence goes beyond keeping systems running. It tests your ability to automate and improve operational processes themselves. This post focuses on advanced CloudWatch features, Systems Manager automation, deployment strategy selection, EventBridge-driven event-based operations, and AWS Config compliance automation.

 

Advanced CloudWatch — Composite Alarms, Metric Math, Anomaly Detection

CloudWatch offers capabilities well beyond simple threshold alarms.

Composite Alarms combine multiple alarms using AND/OR logic. For example, "CPU is high AND network traffic is low" can be expressed as a single alarm. This reduces false positives from individual alarms and delivers notifications only when a state change is genuinely meaningful.

Metric Math combines multiple CloudWatch metrics using mathematical functions to create derived metrics. For example, you can define ErrorRate as (Errors / Requests) * 100 and use that calculated value as an alarm threshold.

Anomaly Detection uses ML algorithms to learn historical patterns and automatically compute normal ranges. Using a dynamic baseline instead of a fixed threshold is especially effective for metrics with different weekday and weekend patterns.

CloudWatch Logs Insights provides serverless SQL queries for real-time log analysis. You can quickly extract specific error patterns, request latency distributions, and user behavior patterns from large log volumes.

CloudWatch Agent collects custom metrics and logs from inside EC2 instances. When an instance terminates, logs on local disk can be lost. Streaming in real time to CloudWatch Logs via the CloudWatch Agent prevents log loss before instance termination.

Service Quotas integrates with CloudWatch, automatically publishing usage metrics to the AWS/Usage namespace. As you approach quota limits, CloudWatch alarms provide advance notification and can trigger automatic quota increase requests.

 

Deployment Strategy Comparison — All Four Strategies Side by Side

| Strategy | Downtime | Rollback Speed | Cost | Risk | Primary Use Case | |----------|---------|----------------|------|------|-----------------| | All-at-once | Yes | Redeploy required | Low | High | Dev/test environments | | Rolling | None | Redeploy required | Low | Medium | Cost-first priority | | Canary | None | Fast | Medium | Low | Progressive validation needed | | Blue/Green | None | Immediate | High | Very low | Production, regulated environments |

Rolling deployment updates instances sequentially. During the deployment, old and new versions run simultaneously — a mixed-version state. Be careful when deploying alongside database schema changes.

Canary deployment routes a small fraction of traffic to the new version first for validation with real users. If a problem is found, only that small fraction of traffic is affected, minimizing risk.

Blue/Green keeps the old version (Blue) running while the new version (Green) is fully prepared, then switches traffic. Rollback is immediate and complete. The cost doubles because two environments run in parallel, but stability is correspondingly higher.

 

Systems Manager — Automation, Patch Manager, Run Command, Parameter Store

AWS Systems Manager is the operational hub for unified management of EC2 and on-premises servers.

Automation Runbooks define a sequence of operational tasks as code. They automate repetitive operational work like restarting EC2 instances, creating AMIs, and applying patches. Integrating with EventBridge lets you automatically trigger Runbooks when specific events occur.

Patch Manager automatically applies OS security patches to EC2 and on-premises servers. Patch Baselines define approved patches, and Maintenance Windows schedule the patch timing. Hybrid Activation brings on-premises servers under the same management model.

Run Command executes commands remotely on EC2 instances and on-premises servers without SSH. Session Manager provides shell access to instances using IAM-based authentication without SSH keys or open ports. All sessions are logged to CloudTrail.

Parameter Store securely stores configuration data and secret values. It supports KMS encryption and version history. The key difference from Secrets Manager is automatic rotation. Secrets Manager supports automatic rotation for RDS, Redshift, and DocumentDB credentials.

 

EventBridge — Event-Driven Automation and Cross-Account

EventBridge receives events from AWS services, custom applications, and SaaS partner events, then routes them to specified targets. It serves as the central hub for event-driven automation.

Enabling cross-account event buses lets you aggregate events from multiple accounts into a central account for unified monitoring and automation. For example, you can collect EC2 state-change events from many accounts into a central security account and automatically trigger Lambda or Systems Manager Automation.

AWS X-Ray provides distributed tracing that visualizes request flow across microservices. The service map shows the entire architecture at a glance, and you can identify latency and errors at each segment. Combined with CloudWatch Agent, you achieve end-to-end observability.

 

AWS Config — Compliance and Auto-Remediation

AWS Config continuously records configuration changes for AWS resources and evaluates compliance. Config Rules define the desired state, and non-compliant resources are detected automatically.

Auto-Remediation links a Config Rule with a Systems Manager Automation Runbook to automatically fix non-compliant states. For example, if an S3 bucket becomes publicly accessible, access is automatically blocked.

KCL (Kinesis Client Library) checkpoint independence is also important from an operational excellence perspective. KCL stores checkpoints in a DynamoDB table with the same name as the applicationName. When multiple consumer applications process the same stream, each must use a different applicationName so their checkpoints remain independent.

 

Exam Key Points

"Combine multiple alarms with AND/OR logic to reduce false positives" -- CloudWatch Composite Alarm

"Dynamic baseline for anomaly detection, no fixed threshold needed" -- CloudWatch Anomaly Detection

"Prevent log loss before EC2 instance termination" -- CloudWatch Agent real-time streaming

"Advance notification before service quota limit is reached" -- Service Quotas plus CloudWatch integration

"Immediate rollback, regulated environments, higher cost" -- Blue/Green deployment

"Cost-first priority, watch for version mismatch during deployment" -- Rolling deployment

"IAM-based instance access without SSH, session auditing" -- Session Manager

"OS patch automation unified across hybrid environments" -- Patch Manager plus Hybrid Activation

"Automatic secret rotation supported" -- Secrets Manager (Parameter Store does not support rotation)

"Auto-trigger Runbook when a specific event occurs" -- EventBridge plus Systems Manager Automation

"Automatically detect and fix non-compliant resources" -- AWS Config plus Auto-Remediation

"Visualize request flow across distributed microservices" -- AWS X-Ray service map

"Keep KCL checkpoints independent across multiple consumers" -- set different applicationName for each consumer

Back to blog list