In any data pipeline, two fundamental questions always arise: "Who is trying to access this?" and "What are they allowed to do?" These are the twin pillars of cloud security — authentication and authorization. Think of it like entering a secure office building: swiping your badge at the door is authentication (you prove who you are), and the fact that your badge only unlocks certain floors is authorization (it defines what you can access). For the DEA-C01 exam, mastering these two concepts is essential, because almost every security question traces back to one or both of them.
Authentication — Proving Who You Are
In data pipelines, services constantly need to authenticate to other services. Glue needs to read from S3. Lambda needs to write to DynamoDB. EMR needs to process files in S3. None of this can happen without identity verification. And in most cases, it is not a human logging in — it is one AWS service proving its identity to another.
IAM Roles — The Delegation System for Services
IAM (Identity and Access Management) is AWS's central identity system. While IAM users are accounts for humans to log into the console, IAM roles are temporary identities that AWS services borrow to act on your behalf. Think of a role as a signed authorization letter: "This Glue job is authorized to access S3 bucket X on behalf of our team."
When a Glue ETL job starts, it assumes an IAM role. That role declares "I am allowed to read from S3, write logs to CloudWatch, and access the Glue Data Catalog." Without the role, Glue cannot do anything at all — AWS denies all access by default.
Common IAM role patterns in data engineering:
| Scenario | Role Type | Purpose | |----------|-----------|---------| | Glue ETL reads/writes S3 | Glue service role | Glue assumes this to access S3 and the Glue Catalog | | Lambda writes to DynamoDB | Lambda execution role | Lambda uses this to interact with DynamoDB | | EMR processes data from S3 | EC2 instance profile | Attached to every EC2 node in the cluster | | Step Functions calls Lambda and Glue | Step Functions execution role | Orchestrator needs permission to trigger child services |
The most important takeaway: every service-to-service interaction in AWS requires an IAM role. If a pipeline fails with "Access Denied," the first thing to investigate is whether the IAM role exists, is attached correctly, and has the right permissions for the specific resource.
VPC Network Security — Controlling Which Roads Data Travels
Authentication answers "who are you," but network security answers "which road must you take to get here." A VPC (Virtual Private Cloud) is your own private network within AWS — an isolated space separated from the public internet.
By default, when an EC2 instance or Lambda function talks to S3, that traffic travels out to the internet and back. This exposes your data transfer to potential interception and also incurs data transfer costs. VPC Endpoints fix this by routing the traffic entirely through AWS's internal backbone network.
Two types of VPC endpoints matter most for data engineering:
Gateway Endpoints are free and serve S3 and DynamoDB. You add a single route entry to your route table, and all traffic to S3 or DynamoDB from within that VPC automatically flows through the internal AWS network — no internet required.
Interface Endpoints (PrivateLink) work for most other services including Glue, KMS, Secrets Manager, Kinesis, SQS, and more. They provision a private IP address inside your VPC that connects directly to the AWS service. There is a small per-hour cost, but the security benefit is significant.
Security Groups work as virtual firewalls around compute resources like EC2, EMR clusters, and Redshift clusters. For a Redshift cluster, a typical security group configuration allows only port 5439 inbound from the analytics team's VPC CIDR, and only outbound to the S3 VPC endpoint. Everything else is blocked.
AWS PrivateLink extends this concept to cross-account and cross-VPC scenarios. If your data platform account needs to expose a streaming service to a consumer account, PrivateLink creates a private link between them — no internet, no VPN, no exposure.
Credential Management — Never Put Passwords in Your Code
Data pipelines frequently need to connect to external databases, third-party APIs, or on-premises systems — all of which require passwords, connection strings, or API keys. Hardcoding these credentials in source code is a serious security mistake. They end up in Git history, log files, or container images where anyone with access can read them.
AWS Secrets Manager is the right solution for sensitive credentials that need to rotate regularly. Its most important feature is automatic rotation. Suppose your security policy requires all database passwords to be changed every 90 days. Secrets Manager handles this entirely on its own: it invokes a Lambda function that changes the password in the database (RDS, Redshift, DocumentDB), updates the stored secret, and your pipeline automatically retrieves the fresh credentials next time it connects. No manual work, no downtime.
AWS Systems Manager Parameter Store is a lighter-weight alternative for configurations that do not need automatic rotation. It stores both plaintext values (like environment names, feature flags, S3 bucket names) and encrypted values (like API keys) using the SecureString type powered by KMS.
| Feature | Secrets Manager | Parameter Store | |---------|----------------|----------------| | Automatic rotation | Yes, via Lambda integration | No | | Cost | $0.40 per secret per month | Free for standard parameters | | Best use case | DB passwords, API keys, certificates | App config, environment settings | | Cross-account access | Yes | Yes | | Versioning | Yes | Yes |
!Authentication versus authorization
Authorization — Deciding What You Are Allowed to Do
Once identity is confirmed, authorization determines which actions that identity is permitted to take. AWS implements authorization through IAM policies and service-specific access controls — and in data engineering, several layers often stack on top of each other.
IAM Policies — The Written Rules of Permission
An IAM policy is a JSON document that explicitly states "allow or deny these actions on these resources." Policies attach to roles, users, or groups, and AWS evaluates them every time an API call is made.
Three policy patterns appear constantly in data engineering:
Service role policies attach to the IAM roles that AWS services assume. A Glue service role policy typically grants s3:GetObject and s3:PutObject on specific S3 buckets, logs:CreateLogGroup and logs:PutLogEvents for CloudWatch, and glue:GetTable for catalog access.
Resource policies attach directly to a resource rather than an identity. An S3 bucket policy is the most common example — it says "this bucket permits reads only from role ARN X" or "this bucket allows access from AWS account 123456789012." Resource policies are critical when you need cross-account access, because you must explicitly permit the external account on the resource side.
The least privilege principle is the single most important security practice: grant only the permissions that are absolutely necessary, to the narrowest scope possible. Instead of granting s3:GetObject on all buckets (arn:aws:s3:::), scope it to the specific bucket and prefix the job actually reads (arn:aws:s3:::my-data-bucket/raw/). This minimizes exposure if credentials are ever compromised.
Lake Formation Permissions — Fine-Grained Data Lake Access Control
AWS Lake Formation provides centralized governance for your data lake. It sits on top of S3 and the Glue Data Catalog and adds column-level, row-level, and table-level access control — things that plain IAM policies cannot do on their own.
Before Lake Formation existed, controlling access to data lake tables was painful. You had to configure IAM policies for the user, S3 bucket policies for the storage, and Glue Catalog resource policies separately. One inconsistency between them and a user could access data they should not. Lake Formation consolidates all of this into a single permission model.
Permission levels Lake Formation supports:
Database level: grant a team access to all tables within a specific Glue database Table level: restrict a user to only specific tables within a database Column level: hide sensitive columns like credit_card_number or ssn from specific users or groups Row level: use data filters to show only rows matching a condition — for example, a regional sales team sees only rows where region = 'us-west'
Data filters are particularly powerful. The underlying S3 data is stored in one place, but Lake Formation applies filters at query time, so different users see different subsets of the same table without any data duplication.
RBAC (Role-Based Access Control) grants permissions based on a person's organizational role. You define roles like "data analyst" or "data engineer" and attach permissions to those roles. When a new employee joins, you assign them the appropriate role and they inherit all its permissions automatically. This is simple to manage for large teams.
TBAC (Tag-Based Access Control) uses metadata tags to automatically match resources to principals. You tag a table with sensitivity=high and tag a user's IAM principal with clearance=high. Lake Formation sees the matching tags and grants access without you needing to write individual permission entries for every table-user combination. When you have hundreds of tables and dozens of teams, TBAC dramatically reduces the management overhead compared to manually granting permissions one by one.
Redshift Access Control — The Data Warehouse Security Layer
Redshift operates its own internal access control system in addition to IAM, making it a multi-layer security environment.
Redshift users and groups work similarly to PostgreSQL. Inside the database, you create use