A data pipeline is ultimately code. The Lambda function that collects data, the Glue job that transforms it, the S3 bucket that stores it — all of it must be defined and managed as code. Just as software developers use version control and automated deployment for application code, data engineers manage their infrastructure and pipelines as code. This is made possible by IaC (Infrastructure as Code) and CI/CD pipelines. The AWS DEA-C01 exam tests your understanding of these tools and when to use each one correctly.
IaC Tool Comparison — Ways to Define Infrastructure as Code
IaC means "defining infrastructure as code so it can be automatically created, managed, and deleted." Instead of manually creating an S3 bucket in the AWS console, configuring RDS, and deploying Lambda by hand, a single code file does all of that in one step.
| Tool | Language | Characteristics | |------|----------|----------------| | AWS CloudFormation | YAML or JSON | AWS-native IaC. Declarative templates. Supports all AWS resources. | | AWS CDK | Python, TypeScript, Java, etc. | Define infrastructure with programming languages. Compiles to CloudFormation internally. | | AWS SAM | YAML (CloudFormation extension) | Specialized for serverless applications. Simplified definition for Lambda, API Gateway, DynamoDB. |
CloudFormation — Declarative Templates
The idea is "declare the final state you want, and AWS creates it for you." Like architectural blueprints, a CloudFormation template says "build me this infrastructure," and AWS automatically creates the required resources.
A Stack is the core unit of CloudFormation. Related AWS resources are grouped into one stack and can be created, updated, and deleted together.
AWS CDK — Define Infrastructure with Programming Languages
CDK lets you define infrastructure using Python, TypeScript, or other programming languages. It is far more flexible than YAML. You can use loops, conditionals, functions, and classes to dynamically define infrastructure.
For example, if you need to create the same S3 bucket in 10 regions, CloudFormation would require 10 separate templates. CDK solves it with a single loop. CDK code is ultimately synthesized (compiled) into CloudFormation templates before deployment.
AWS SAM — Serverless-Specialized IaC
SAM (Serverless Application Model) is an extension of CloudFormation. It defines serverless resources like Lambda, API Gateway, and DynamoDB with much less boilerplate. The SAM CLI also supports local testing and streamlined deployment.
Lambda — Understanding It from a Data Engineering Perspective
AWS Lambda runs code without any servers. Code only executes when an event triggers it, and you only pay for the duration of execution. In data engineering, Lambda acts as "small automation components" throughout a pipeline.
Key Lambda constraints you must know
Concurrency: The number of Lambda instances running simultaneously. The default account-level limit is 1,000.
Reserved Concurrency: Reserves a specific number of concurrent executions for a particular Lambda. Prevents other Lambdas from consuming the entire limit. Provisioned Concurrency: Pre-warms instances so they are ready to respond instantly — eliminating cold starts (the delay that occurs when a function runs for the first time after a period of inactivity).
Timeout: Maximum 15 minutes. Any ETL job that needs longer than 15 minutes must be moved to Glue ETL or EMR. Lambda is best suited for short, fast tasks in data engineering.
Memory: Configurable from 128 MB to 10 GB. In Lambda, increasing memory also proportionally increases CPU performance. If data processing is slow, increasing memory is the first optimization to try.
EFS mount: Lambda can mount Amazon EFS (Elastic File System) to access large files. While Lambda's temporary storage (/tmp) has a maximum of 10 GB, EFS is the right choice for files shared across multiple functions or datasets larger than what /tmp can hold.
Runtime: Python is by far the most common Lambda runtime for data engineering. Libraries like pandas, numpy, and boto3 can be added as Lambda Layers.
CI/CD Pipeline — Automated Deployment for Data Pipelines Too
CI/CD (Continuous Integration / Continuous Delivery) is a system where code changes are automatically tested and deployed. Because data pipelines are code, the same CI/CD principles that apply to software development apply here too.
A typical AWS CI/CD pipeline flow:
CodeCommit: AWS-managed Git repository, similar to GitHub. CodeBuild: Build and test automation, runs in containers. CodeDeploy: Automates application code deployment to EC2, Lambda, and ECS. CodePipeline: Connects all of the above stages into a single pipeline.
Data pipeline CI/CD considerations: Glue ETL script change → automated tests → deploy to S3 CloudFormation/CDK template change → automated stack update Lambda code change → automated deployment + blue/green deployment for safe rollback
Core Distributed Computing Concepts
To work with distributed processing systems like Spark and EMR, you need to understand a few fundamental concepts. These appear frequently in DEA-C01 exam questions about performance issues and system design.
Shuffling
The process of redistributing data across nodes. For example, a GROUP BY country operation requires all records for the same country to be on the same node. Moving data across the network to accomplish this is shuffling.
Shuffling is expensive because data travels across the network between nodes. It is one of the leading causes of Spark performance problems. Minimizing shuffling is a central theme in Spark optimization.
Partitioning
Dividing data into logical units for parallel processing. If you split 100 million records into 100 partitions, 100 cores can each process one partition simultaneously. Too few partitions means low parallelism; too many partitions increases overhead from coordination.
Data Skew
When data is distributed extremely unevenly across partitions. For example, if one partition out of 100 holds 80% of the total data, the node processing that partition becomes a bottleneck — all other nodes finish and then wait for the single overloaded node. The entire job runs at the speed of the slowest partition.
How to fix skew: Change the partition key to a column with more even distribution Artificially split the skewed partition into smaller pieces (a technique called salting)
Serialization
To send data between nodes, in-memory objects must be converted to a byte stream (serialization), and the receiving node converts them back to objects (deserialization). Kryo serialization is faster and more compact than Java's default serialization and is the recommended choice for Spark jobs.
Exam Key Points Summary
| Keyword | Tool/Concept | |---------|-------------| | YAML/JSON declarative infrastructure definition | AWS CloudFormation | | Define infrastructure with a programming language (Python/TypeScript) | AWS CDK | | Serverless-specialized IaC (Lambda, API Gateway) | AWS SAM | | Access large shared files from Lambda | EFS mount | | Lambda job exceeds 15 minutes | Switch to Glue ETL or EMR | | Eliminate Lambda cold starts | Provisioned Concurrency | | Automated build and deployment on code change | CodePipeline (CodeBuild + CodeDeploy) | | Data redistribution across nodes (expensive Spark operation) | Shuffling | | Data concentrated in specific partitions | Data Skew | | Fix uneven partition distribution | Change partition key or use Salting |
CDK does not replace CloudFormation — it is an abstraction layer built on top of CloudFormation. CDK code always compiles down to a CloudFormation template before deployment.