Data transformation is like a factory production line that turns raw materials into finished goods. Raw data collected from various sources is often not ready for analysis — it may need cleaning, reformatting, enriching with other data, or converting to a more efficient structure. ETL (Extract, Transform, Load) is the term for this entire processing pipeline. In the AWS DEA-C01 exam, questions focus heavily on "which transformation tool to pick, which file format to use, and how to design distributed processing."
ETL Service Comparison — Which Tool Fits Which Situation?
AWS offers multiple services for data transformation. Understanding what each one is good at is the key to answering exam questions correctly.
| Service | Characteristics | Best situation | |---------|----------------|----------------| | AWS Glue ETL | Fully serverless, Apache Spark-based, auto-scaling | Batch transformation of structured/semi-structured data, moving data between S3/Redshift/RDS | | Amazon EMR | Manage your own Hadoop/Spark cluster | Large-scale custom processing, complex pipelines, cost optimization with Spot instances | | AWS Lambda | Fully serverless, event-driven, max 15 minutes | Simple transformations, small data, event-triggered jobs | | Amazon Redshift | SQL-based transformation | Transforming and aggregating data already loaded in Redshift |
Decision rules: Building a new serverless ETL pipeline? Use Glue ETL. Have an existing Hadoop/Spark workload or need fine-grained cluster control? Use EMR. Simple transformation triggered by a file upload or API call? Use Lambda. Want to transform data that is already inside Redshift? Use Redshift SQL.
AWS Glue ETL — The Serverless Data Processing Factory
AWS Glue ETL is a serverless service that lets you perform large-scale data transformations without managing any infrastructure. Think of it like a chef (Glue) who receives ingredients (raw data), follows a recipe (ETL script) to cook the dish (transform), and then plates it (delivers to S3, Redshift, etc.). You never have to buy or maintain the kitchen.
What is a DynamicFrame?
A DynamicFrame is the basic data unit in Glue ETL, similar to Apache Spark's DataFrame but smarter about handling real-world messy data. If some records have an "age" field and others do not, or if the same field has inconsistent types across rows, DynamicFrame handles this automatically without throwing errors. This makes it ideal for working with the kind of inconsistent, imperfect data that comes from real production systems.
Glue Job Bookmarks
Bookmarks are Glue's way of "remembering where it left off." Imagine new log files land in S3 every night. Without bookmarks, every run would reprocess all files from the beginning. With bookmarks enabled, Glue tracks which files and records it has already processed and skips them on the next run — like a bookmark in a book telling you which page to start from.
Three job types:
Spark jobs: Best for large-scale batch processing. Distributed across multiple nodes, capable of handling terabytes of data. Streaming jobs: Process real-time data from Kinesis or MSK in micro-batches. Python Shell jobs: Simple Python scripts for lightweight tasks, small datasets, or API calls.
Glue Studio
A visual drag-and-drop editor for building ETL pipelines without writing code. You connect data sources, transformation steps, and destinations visually, and Glue Studio generates the PySpark code automatically. This makes ETL pipeline creation accessible to non-developers.
File Format Selection — Which Container Should Your Data Live In?
File format choices have a major impact on performance and cost. Think of it like choosing containers for food: the right container keeps food fresh longer, is easier to transport, and takes up less space. The right data format reads faster, costs less to query, and compresses better.
| Format | Structure | Compression | Use case | |--------|-----------|-------------|----------| | CSV | Row-based, text | Low | Human-readable raw data, small scale | | JSON | Row-based, text | Low | API responses, nested structure data | | Parquet | Column-based (columnar), binary | High | Analytical queries (Athena, Redshift Spectrum, Spark) | | ORC | Column-based, binary | High | Apache Hive workloads, EMR | | Avro | Row-based, binary | Medium | Streaming, frequently changing schemas |
Why columnar formats (Parquet/ORC) are superior for analytics
Imagine looking up a topic in a book. Instead of reading every page, you go to the table of contents and jump directly to the right page. Columnar formats work the same way. Analytical queries typically need only a few columns out of potentially hundreds. Parquet stores data column by column, so only reads the revenue and date columns — skipping all other columns entirely. On tables with hundreds of columns, this dramatically reduces both query time and cost.
Why Avro is strong for streaming
Avro stores the schema alongside the data. When the schema changes — a new field added, an old one removed — you do not need to rewrite existing records. Old records still work correctly with the new schema. This schema evolution capability makes Avro the natural fit for real-time data streams where the structure of incoming data may change over time.
Format selection rules: Analytical queries with Athena or Redshift Spectrum are the primary goal → Parquet Apache Hive or EMR workloads → ORC Streaming data or frequently changing schema → Avro Received CSV files and want to query them with Athena efficiently → Convert to Parquet with Glue first
Amazon EMR — You Drive the Distributed Processing Cluster
Amazon EMR (Elastic MapReduce) is a service that runs distributed processing frameworks like Apache Hadoop and Apache Spark on a cluster of EC2 instances in AWS. If Glue ETL is a "taxi" (serverless, no driving required), EMR is more like a "personal vehicle" — you have complete control but you also manage it yourself.
The three EMR instance group roles
Master Node: The brain of the cluster. It distributes work and coordinates the cluster. There is always exactly one master node, and if it fails the whole cluster fails — so use On-Demand instances here. Core Node: Handles actual data processing and HDFS (Hadoop Distributed File System) storage. If a core node disappears, data stored on it can be lost — use On-Demand instances here too. Task Node: Handles processing only, with no HDFS storage. Since task nodes do not hold data, they can be added or removed at any time without data loss. This makes them perfect for Spot Instances (unused AWS capacity available at a discount), significantly reducing costs.
Spot Instance strategy:
| Node type | Recommended instance | Reason | |-----------|---------------------|--------| | Master | On-Demand | Essential for cluster stability | | Core | On-Demand | Prevent data loss | | Task | Spot Instances | No data storage, great for cost savings |
EMR Serverless
A way to run Spark or Hive jobs without configuring a cluster at all. Similar to Glue ETL in the serverless respect, but the key difference is that EMR Serverless runs pure Spark/Hive code as-is, while Glue ETL provides its own DynamicFrame API layer.
Exam Key Points Summary
| Keyword | Service/Concept | |---------|----------------| | Serverless ETL, Spark-based, auto-scaling | AWS Glue ETL | | Large-scale custom Hadoop/Spark processing | Amazon EMR | | Track already-processed data, prevent reprocessing | Glue Bookmarks | | Auto-handle schema mismatches | Glue DynamicFrame | | Build ETL pipeline without code | Glue Studio | | Columnar format optimized for analytical queries | Parquet | | Columnar format for Hive/EMR workloads | ORC | | Streaming data, frequently changing schema | Avro | | EMR cost reduction | Use Spot Instances on Task Nodes | | ETL job exceeding Lambda's 15-minute limit | Switch to Glue or EMR |
Converting CSV to Parquet is a recurring DEA-C01 exam topic because it dramatically improves both the speed and cost of Athena queries.