Incident and Event Response

Master the core of the DOP-C02 Incident domain (14%): EventBridge-driven automation, ASG Lifecycle Hook patterns, and CI/CD pipeline troubleshooting.

The Incident and Event Response domain makes up 14% of the DOP-C02 exam. The questions go beyond simple tool knowledge, asking you to design complete event-driven automation chains and identify the right event source for each scenario.

 

Event Sources and Processing Patterns

Understanding event sources precisely is critical for a DevOps engineer. AWS provides multiple event-generating services, each with a specific purpose.

AWS Health + EventBridge Pattern

AWS Health publishes real-time events about service disruptions, planned maintenance, and instance retirement affecting your account. Using in EventBridge rules, you can immediately capture Health events.

Consider this real-world scenario: a company running thousands of EC2 instances receives retirement notices from AWS. Manual console monitoring leads to delayed responses and service degradation. The solution is an EventBridge rule that detects and triggers an SSM Automation runbook to automatically stop and restart affected instances.

| Event Source | source Value | Representative Event Types | |-------------|-------------|--------------------------| | AWS Health | aws.health | EC2 retirement, service disruption, maintenance | | CloudTrail | (natively integrated) | All AWS API calls | | EC2 Auto Scaling | aws.autoscaling | Instance launch/terminate lifecycle events | | CodePipeline | aws.codepipeline | Pipeline state changes |

Do not confuse Health with CloudTrail. AWS infrastructure state changes come from Health; API calls made by users come from CloudTrail.

 

CloudTrail + EventBridge — Real-Time API Event Surveillance

When you need near-real-time detection of sensitive API calls like IAM user creation or security group modifications, leverage the native integration between CloudTrail and EventBridge.

Financial institution scenario: IAM policy prohibits direct IAM user creation, but some developers bypass this. The response architecture works as follows.

Set EventBridge rule pattern: Register Lambda as the target Lambda calls + to immediately disable the user SNS sends notification to the security team

This architecture completes detect→disable→notify within seconds of the API call. Periodically querying CloudTrail logs with Athena introduces tens of minutes of delay and is unsuitable for real-time response.

What about SSH port 22 being opened to 0.0.0.0/0 on a security group? The same pattern applies: detect events with EventBridge conditions filtering , then auto-trigger Lambda to call .

 

SQS Dead Letter Queue — Isolating Poison Messages

In large-scale event processing systems, certain messages that repeatedly fail to process and block the normal message flow are called "poison pill" messages.

SQS Dead Letter Queue (DLQ) automatically moves messages that exceed from the source queue to a separate DLQ. The key benefit is isolating poison messages to protect normal flow, while DLQ-isolated messages are preserved in their original form for later root cause analysis and reprocessing.

When Lambda processes messages via SQS event source mapping, all messages in a batch must succeed before an ACK is sent. On failure, the entire batch is retried. Once is reached, the message moves to the DLQ. The SQS console's DLQ Redrive feature lets you send messages back to the source queue after analysis.

 

Kinesis Data Streams Enhanced Fan-Out — Scaling Consumers

When multiple Lambda consumers simultaneously try to read from the same DynamoDB Streams shard, errors occur. DynamoDB Streams allows at most 2 concurrent consumers per shard.

Switching to Kinesis Data Streams with Enhanced Fan-Out gives each consumer a dedicated 2 MB/s throughput per shard. Consumers operate independently with no contention. DynamoDB also supports direct integration with Kinesis, enabling seamless transition without code changes. Compare this with SNS + SQS fan-out, which adds a message broker layer, whereas Kinesis Enhanced Fan-Out directly scales stream consumers.

 

Event-Driven Configuration Changes and Auto-Recovery

CloudWatch Alarm + SSM Run Command — Process Auto-Recovery

Imagine a game server application where the instance itself is healthy but the game server process crashes intermittently. Auto Scaling health checks only detect instance-level failures and cannot detect application process failures.

The solution is to use the CloudWatch Agent's plugin to collect the running state of a specific process as a custom metric. When the process count drops to zero, a CloudWatch Alarm fires, and the alarm action triggers SSM Run Command to restart only the process on the affected instance.

The advantage of this approach: no instance replacement means no player session disconnections. The Auto Scaling DesiredCapacity remains unchanged. Commands execute remotely through SSM Agent without SSH.

 

ASG Lifecycle Hook — Log Collection Before Termination and Instance Preservation

When an Auto Scaling group terminates an instance, the logs on that instance disappear with it. ASG Lifecycle Hooks are the key tool when you need to preserve logs for failure root cause analysis.

A Lifecycle Hook inserts a stage when an instance transitions from to . During this wait period (default 3600 seconds, maximum 48 hours), the instance is not actually terminated and remains accessible.

Two critical usage patterns appear on the exam.

Pattern 1, log collection: EventBridge detects the event, triggers Lambda to send logs to S3, then calls to allow final termination. Termination only proceeds after log transfer completes.

Pattern 2, service discovery synchronization: Hooks on both launch () and termination () trigger Lambda to automatically synchronize external registries.

Common trap: EventBridge + Run Command looks conceptually similar, but SSM commands cannot reach an already-terminating instance. You must pause the termination with a Lifecycle Hook to make command execution possible.

 

CI/CD Pipeline Troubleshooting

When CodeCommit → CodePipeline Auto-Trigger Does Not Fire

A developer pushes code to the main branch but the pipeline does not start automatically. Manual execution works fine. What should you check first?

The answer is to check the EventBridge rule status. When CodePipeline configures a CodeCommit source, it automatically creates an EventBridge rule that detects events. If this rule is DISABLED or points to the wrong pipeline ARN, the auto-trigger fails. The fact that manual execution works proves there is no IAM permission issue, pointing to the trigger mechanism itself.

When All CodeDeploy Deployment Events Show Skipped Status

The deployment does not fail; it never starts and shows Skipped status. CodeDeploy Agent operates in pull mode. The agent periodically polls the CodeDeploy service endpoint to check for deployment commands.

In a private subnet with neither a NAT Gateway nor VPC endpoints, the agent cannot reach and polling fails. All events showing Skipped means the agent never received the command at all. Distinguish between Skipped and Failed: Failed means the agent received and executed the command but it failed. Skipped means the agent never received the command.

Solution: create a VPC interface endpoint, or provide NAT Gateway-based internet access.

X-Ray and Step Functions for Tracing Complex Workflows

Implementing a multi-step security incident automation workflow with Lambda chaining makes it very hard to track failures in individual steps. AWS Step Functions Standard Workflow preserves each state's input/output and transition history for 90 days. Retry/Catch lets you declaratively define per-step retry logic. The pattern of EventBridge detecting an AWS Health event and triggering a Step Functions workflow is the AWS standard architecture for security incident automation.

 

Exam Key Points

"Auto-detect EC2 retirement + auto-restart" -- EventBridge(source: aws.health, AWS_EC2_INSTANCE_RETIREMENT_SCHEDULED) + SSM Automation

"Detect IAM user creation instantly + auto-disable" -- CloudTrail + EventBridge(source: aws.iam, CreateUser) + Lambda

"Isolate repeatedly failing messages + protect normal flow" -- SQS Dead Letter Queue(maxReceiveCount)

"Scale DynamoDB Streams consumers + no throttling" -- Kinesis Data Streams Enhanced Fan-Out

"Detect process crash + recover without instance replacement" -- CloudWatch Agent procstat + CloudWatch Alarm + SSM Run Command

"Preserve logs before instance termination" -- ASG Lifecycle Hook(Terminating:Wait) + Lambda + complete-lifecycle-action

"CodeCommit push does not trigger pipeline" -- Check EventBridge rule ENABLED/DISABLED status

"All CodeDeploy events Skipped" -- Private subnet with no VPC endpoint or NAT (Agent cannot poll)

Back to blog list