Amazon EMR Complete Guide (DEA-C01 Exam Essentials)
Introduction
Amazon EMR (Elastic MapReduce) appears across both Domain 1: Data Ingestion and Transformation (34%) and Domain 2: Data Store Management (26%) of the DEA-C01 exam.
EMR is a fully managed service that makes it easy to run big data frameworks like Hadoop and Spark on AWS. The exam frequently tests cluster architecture, storage options, and cluster types.
This guide covers the key concepts you need to know for exam preparation.
---
What is Amazon EMR?
Let's start with the fundamentals of EMR.
| Aspect | Detail | | --- | --- | | Definition | Fully managed big data processing service based on Apache Hadoop | | Supported frameworks | Hadoop, Apache Spark, HBase, Presto, Flink | | Key use cases | Log analysis, financial analysis, ETL jobs | | Infrastructure | Amazon EC2 + Amazon S3 | | Access methods | AWS Console, CLI, SDK, EMR API, SSH |
EMR runs on EC2 and allows direct OS access via SSH. While it is fully managed, the ability to access the underlying OS distinguishes it from other managed services.
---
Cluster Node Architecture (3 Node Types)
EMR clusters consist of three node types based on their roles. This structure appears on the exam regularly.
| Node Type | Role | Data Storage | Computation | Notes | | --- | --- | --- | --- | --- | | Master Node | Coordinates and manages the entire cluster | X | X | Task assignment, failure recovery, cluster state management | | Core Node | Data storage + computation | O | O | Handles data replication, essential for cluster operation | | Task Node | Computation only (optional) | X | O | Can use Spot Instances for cost savings |
Task Nodes do not store data and handle computation only. Scenarios involving Spot Instances on Task Nodes for cost reduction frequently appear on the exam.
!EMR cluster node types: Master, Core, Task
Storage Options Comparison (HDFS vs EMRFS vs Local)
Comparing storage options is a common exam question on DEA-C01.
| Storage | Description | Persistence | Key Use Cases | | --- | --- | --- | --- | | HDFS | Distributed file system, 128MB block splitting and replication | Temporary (deleted when cluster terminates) | Intermediate processing data, temporary storage | | EMRFS | Uses S3 as a Hadoop-compatible file system | Persistent (retained after cluster termination) | Input/output data, long-term retention | | Local File System | EC2 instance local disk | Temporary (deleted when instance terminates) | Caching, temporary processing data |
When the exam presents a scenario about "retaining data after cluster termination," the answer is EMRFS (S3). HDFS data is lost when the cluster terminates.
---
External Metastore
By default, the Hive metastore is stored in a local MySQL database on the Master Node. When the cluster terminates, the metastore is lost along with it.
To solve this, external metastores are used.
| Option | Characteristics | | --- | --- | | AWS Glue Data Catalog | Fully managed, integrates with Athena and Redshift | | Amazon RDS / Aurora | High availability and durability, suitable for large metadata management |
When files are added directly to HDFS or S3, Hive may not recognize new partitions. The MSCK REPAIR TABLE command resolves this.
This command scans the file system for new partitions and syncs them with the Hive metastore.
---
Cluster Types: Transient vs Long-Running
| Type | Transient Cluster | Long-Running Cluster | | --- | --- | --- | | Lifespan | Auto-terminates after job completion | Runs continuously | | Key use cases | Batch ETL, periodic processing jobs | Interactive analysis, always-on services | | Cost | Billed only for execution time, highly cost-efficient | Continuous billing for cluster uptime |
Periodic batch processing uses transient clusters, while always-on access requires long-running clusters.
---
Amazon EMR Serverless
EMR Serverless is a serverless option that lets you run Spark or Hive applications without provisioning or managing clusters.
Key features:
AWS automatically handles cluster capacity optimization and scaling Supports Apache Spark and Apache Hive Focus on analytics workloads without infrastructure management overhead
When the exam presents a scenario about "running Spark jobs without managing clusters," the answer is EMR Serverless.
---
AWS Graviton2 and Spark Memory Overhead
AWS Graviton2 Instances
ARM-based custom processors that deliver up to 40% better price-performance compared to equivalent x86 instances. Particularly effective for data-intensive workloads like Spark and Hadoop.
Spark Memory Overhead
Spark incurs approximately 10% additional memory overhead on driver and executor requested memory. This is used for internal operations like shuffle tasks and task execution. This overhead must be accounted for in memory configuration to prevent OOM (Out-of-Memory) errors.
---
Quick Reference Summary
| Keyword | Key Concept | | --- | --- | | EMR + Spark | Big data in-memory processing | | Master Node | Cluster management only (no storage or computation) | | Core Node | Storage + computation (required) | | Task Node | Computation only (optional, Spot Instance compatible) | | HDFS | Temporary storage (deleted on cluster termination) | | EMRFS | S3-based persistent storage | | MSCK REPAIR TABLE | Hive partition metadata sync | | Transient Cluster | Batch ETL, cost-efficient | | EMR Serverless | No cluster management needed | | Graviton2 | Up to 40% better price-performance vs x86 |
---
Wrap-Up
Amazon EMR is an essential service for big data processing scenarios in DEA-C01.
For the exam, it is particularly important to clearly understand the following:
The role differences among the three node types The persistence difference between HDFS and EMRFS The cost difference between transient and long-running clusters
Mastering these concepts will prepare you to answer most DEA-C01 big data processing questions.