Building a pipeline is not the finish line — you need to keep watching it to know it is running correctly. If a data processing job fails and you find out six hours later, bad data may have already flowed into reports. Monitoring is about detecting problems quickly. Data quality management is about preventing bad data from passing through the pipeline in the first place. DEA-C01 exam questions on these topics appear as operational stability and data reliability scenarios.
CloudWatch — AWS's Central Monitoring Service
CloudWatch is the monitoring platform where you can observe nearly everything happening in AWS. Think of it like a health-check machine that measures service vitals as numbers and sends alerts when something looks wrong.
Metrics: Numerical measurements over time. Lambda function invocation count, Glue job run duration, Kinesis stream latency — these are all metrics. Some metrics are collected automatically; others are custom metrics your code publishes directly.
Alarms: Alarms watch a metric and trigger an action when it crosses a threshold. For example, if the Lambda error rate exceeds 5%, send an email via SNS or trigger an Auto Scaling policy.
Logs: CloudWatch Logs collects text log output from services. Lambda function print statements, Glue job error messages, and EMR system logs all land in CloudWatch Logs.
Logs Insights: Query your CloudWatch logs using an SQL-like language. Search for specific error messages in large log volumes, or aggregate error counts by time window.
Per-Service Monitoring
Each AWS service publishes its own metrics to CloudWatch.
Glue monitoring: : How much data the Glue job read : Number of failed tasks : JVM heap memory usage Job run history is also viewable directly in the Glue console
EMR monitoring: : Whether the cluster is sitting idle with no active work : Proportion of containers waiting because of insufficient resources The YARN ResourceManager UI shows individual job progress
Redshift monitoring: : Current connection count , : Disk read and write latency : CPU usage Query execution plans reveal slow queries and their bottlenecks
Kinesis IteratorAge: The single most important Kinesis metric. IteratorAge is the age in milliseconds of the oldest unprocessed record in the stream. When this number is near zero, consumers are keeping up with producers in real time. When it grows continuously, consumers are falling behind — a clear sign you need more processing capacity.
Performance Troubleshooting
What to check when each service runs slowly.
Glue performance: DPU (Data Processing Unit) is Glue's unit of compute. Too few DPUs means slow processing. Too many wastes money. The right number depends on data volume and transformation complexity. Look at and memory metrics together to find the actual bottleneck.
EMR performance: When a job is slow, check instance types first. Memory-intensive Spark jobs benefit from memory-optimized instances (R series). Compute-intensive jobs benefit from compute-optimized instances (C series). Mixing Spot instances reduces cost but introduces interruption risk.
Kinesis shards: When Kinesis throughput hits its limit, IteratorAge grows. Each shard handles 1 MB/s write and 2 MB/s read. When throughput is insufficient, add more shards through resharding.
Five Dimensions of Data Quality
Data quality is not simply right or wrong. It is evaluated across multiple dimensions.
| Dimension | Meaning | Example | |-----------|---------|----------| | Completeness | Are required values present? | Percentage of records with NULL email | | Uniqueness | Are there no duplicates? | Same order ID appearing twice | | Validity | Do values match expected format and range? | Date of 2099, age of -5 | | Consistency | Do values match across systems? | CRM customer count differs from warehouse | | Timeliness | Does data arrive when expected? | Hourly data arriving 2 hours late |
!5 dimensions of data quality
DataBrew Quality Rules
AWS Glue DataBrew lets you define and automatically enforce data quality rules through a visual interface.
Example rules: Column must be between 0 and 150 Column must not contain NULLs Column must be unique Column must be one of: "pending", "completed", "cancelled"
DataBrew applies these rules to the full dataset automatically and reports the percentage of records that violate each rule, along with sample violating rows. If a rule violation rate exceeds a threshold, the job can be marked as failed or a warning raised.
Sampling Techniques
Inspecting hundreds of millions of records takes too much time and money. Sampling checks a subset and uses it to estimate the quality of the full dataset.
Random Sampling: Select records at random from the full dataset. Simple to implement, unbiased results. Works best when data is uniformly distributed.
Stratified Sampling: Divide data into groups (strata) and sample a fixed proportion from each group. For example, when sampling order data by region, ensure each region is represented proportionally. More accurate than random sampling when group sizes differ dramatically.
Systematic Sampling: Select every Nth record. For example, pick every 100th record from a 1-million-record dataset. Useful when data is sorted chronologically and you want to detect periodic patterns.
Data Skew
Data skew is when data is unevenly distributed across nodes or partitions in a distributed processing system. Most nodes finish quickly, but the whole job waits for the one overloaded node.
Example: When aggregating order data by customer ID in Spark, if one large enterprise customer has millions of orders while all other customers have dozens, that one partition is overwhelmed.
Skew solutions: Salting: Add a random suffix to skewed keys to spread them across multiple partitions Repartitioning: Adjust partition count using or Broadcast joins: Replicate a small table to every node to eliminate the expensive shuffle join
Exam Key Points
"Collect AWS service metrics and set alarms" → CloudWatch "Query specific errors in logs" → CloudWatch Logs Insights "Detect Kinesis consumer lag" → IteratorAge metric "Glue job processing is slow" → Adjust DPU count "EMR running out of memory" → Switch to memory-optimized instances "Kinesis throughput insufficient" → Add more shards "Check data for NULLs and out-of-range values" → DataBrew quality rules "Too much data to inspect fully" → Use sampling "Some nodes overloaded in distributed processing" → Data skew