AWS Glue Complete Guide (DEA-C01 Exam Essentials)
Introduction
AWS Glue is the most heavily tested service in the Domain 1: Data Ingestion and Transformation (34%) section of the DEA-C01 exam.
AWS Glue is a fully managed serverless ETL service that automates the entire process of extracting, transforming, and loading data. The exam frequently presents scenario-based questions that require you to choose the right Glue component or feature.
This guide covers the key concepts you need to know for exam preparation.
---
What is AWS Glue?
Let's start with the fundamentals of AWS Glue.
| Aspect | Detail | | --- | --- | | Definition | Fully managed serverless ETL service | | Processing engine | Apache Spark (distributed computing) | | ETL code | Auto-generated and customizable in Python (PySpark) or Scala | | Infrastructure management | Not required (serverless) | | Billing unit | DPU (Data Processing Unit) | | Key integrations | S3, Redshift, RDS, Athena, EMR |
On the exam, when you see a scenario requiring "serverless ETL", the answer is AWS Glue.
---
DPU (Data Processing Unit)
DPU is the computing resource unit in AWS Glue.
| Aspect | Detail | | --- | --- | | 1 DPU | 4 vCPU + 16 GB memory | | Billing | DPU count x processing time | | Scaling | Auto-adjusted (manual specification also available) | | Monitoring | Check optimal DPU in the Job Run Monitoring section of the AWS Glue console |
If DPU is insufficient, you can increase it by raising the parameter.
---
DynamicFrames vs Spark DataFrames
There are two frameworks for processing data in Glue. Understanding their differences is important.
| Comparison | Spark DataFrame | AWS Glue DynamicFrames | | --- | --- | --- | | Schema | Must be predefined | Automatically handled (no predefinition needed) | | Semi-structured data | Difficult to process | Automatically handles JSON, XML, CSV, etc. | | Nested structures | Manual flattening required | Automatically handles nested structures and arrays | | Best for | Structured data | Data lakes with flexible schemas |
On the exam, DynamicFrames is often the correct answer when dealing with data that has frequently changing or unknown schemas.
---
Preventing Duplicate Processing with Job Bookmarks
Job Bookmarks is a feature that tracks processing state across ETL job runs.
Key characteristics:
Already-processed data is not reprocessed (incremental loading) Prevents unnecessary reprocessing in large datasets, saving cost and time Saves checkpoints of processed data for reference in subsequent runs
When the exam presents a scenario about "processing only new data", Job Bookmarks is the answer.
---
AWS Glue Crawlers
Crawlers automatically scan data sources to discover schemas and keep the Glue Data Catalog up to date.
| Feature | Detail | | --- | --- | | Auto schema discovery | Scans data sources like S3, RDS and auto-creates/updates tables | | Supported formats | CSV, JSON, Parquet, ORC, etc. | | Partition detection | Automatically detects partitions to improve query performance | | Scheduling | Regular execution to keep Data Catalog current | | Access control | Fine-grained metadata access control via IAM policies |
The key exam point here is understanding the two layers of access control.
These two layers are managed separately.
---
Schema Registry for Streaming Data
Schema Registry ensures schema consistency in streaming data from Kinesis and Kafka.
Key features:
Schema version control Schema validation when producers add new records to streams Record rejection or transformation on schema mismatch Integration with KPL (Kinesis Producer Library) and KCL (Kinesis Client Library)
When the exam asks about "ensuring schema consistency in Kinesis/Kafka streaming data", the answer is Glue Schema Registry.
---
Key Transformation Features
Glue provides several important transformation capabilities.
| Feature | Description | | --- | --- | | ResolveChoice | Handles mixed data types in a column (e.g., strings and integers in the same column) | | FindMatches ML | Identifies matching entity records without common keys (ML-based) | | Pivot | Converts rows to columns (optimizes analytical queries) | | CTAS | Creates new tables from SELECT results (format conversion, summary tables) |
When the exam presents a scenario about "detecting and merging duplicate records using ML", the answer is FindMatches ML Transform.
---
Execution Environment Comparison
Glue offers different execution environments depending on the workload type.
| Environment | Characteristics | Best For | | --- | --- | --- | | Spark | Distributed computing, large-scale processing | Large ETL jobs, complex transformations | | Python Shell | Lightweight script execution | Simple transformations, AWS service integration | | Ray | Python parallel processing | ML/AI workloads, Python scale-out | | Streaming ETL | Real-time stream processing | Kinesis and Kafka real-time data processing |
Key selection points that frequently appear on the exam:
Lightweight Python ETL → Python Shell ML workload Python scale-out → Ray
---
Performance Optimization Strategies
Key strategies for improving ETL job performance.
| Strategy | Description | | --- | --- | | Data partitioning | Partition by date, region, etc. to improve query performance | | Partition indexes | GetPartitions API queries only matching partitions, improving large-partition performance | | Broadcast Join | Copies small datasets to all nodes, minimizing shuffling in large joins | | Small file consolidation | Many small files degrade Spark performance; consolidate into larger files | | Server-side filtering | Pre-filter data using partition predicates when creating DynamicFrames |
---
AWS Glue DataBrew
DataBrew is a visual data preparation tool that lets you clean and normalize data without writing code.
| Feature | Detail | | --- | --- | | No-code transformations | 250+ pre-built transformations (drag and drop) | | Data profiling | Per-column statistics, automatic detection of missing values, duplicates, and outliers | | Custom data quality rules | Define validation rules based on business requirements | | PII detection and masking | ML-based detection of personal information with masking and anonymization | | Serverless | No infrastructure management required |
Common exam scenarios:
Need data cleaning and profiling without coding → AWS Glue DataBrew Need PII detection and masking → DataBrew or Glue Sensitive Data Detection
---