Encryption is the digital equivalent of a padlock. It transforms data into an unreadable form so that without the right key, no one can understand the contents — even if they physically obtain the storage device or intercept network traffic. In cloud data engineering, encryption is not optional. Whenever you handle customer personal data, financial records, or medical information, encryption is both a security requirement and, in many cases, a legal obligation. This guide explains how AWS implements encryption, how to mask sensitive data, and how to meet compliance requirements — starting from the basics.
Encryption at Rest
Encryption at rest means data is encrypted while it sits on a disk or storage medium. Even if someone physically steals a hard drive or gains unauthorized access to a storage system, they cannot read the data without the decryption key.
Amazon S3 Encryption Options
S3 offers three server-side encryption (SSE) methods, each suited to different security requirements.
SSE-S3 is S3's default encryption method using AWS-managed keys. It is enabled by default on all new S3 buckets at no extra cost. AWS manages the entire key lifecycle — creation, rotation, and storage. The tradeoff is that you have no visibility into key usage. You cannot see who decrypted which object using which key. For general-purpose data with no strict compliance requirements, SSE-S3 is perfectly adequate.
SSE-KMS uses a customer-managed key (CMK) from AWS Key Management Service. You create the KMS key yourself, define who can use it through key policies, and retain control over key rotation. The most significant advantage is auditability: every time anyone uses the key to encrypt or decrypt an S3 object, that event is recorded in CloudTrail. For financial data, healthcare records, or any regulated data, SSE-KMS is the standard choice. There is a cost per KMS API request, but the audit trail it provides is worth it for compliance.
SSE-C means the customer provides the encryption key directly. When uploading an object, you pass the key in an HTTP header. S3 uses that key to encrypt the data and then discards the key — it never stores it. When downloading, you must provide the same key again. This gives you complete control over your keys, but also means you are entirely responsible for key management and secure storage. The operational complexity is high, so SSE-C is used only in scenarios with the strictest key custody requirements.
| Method | Key Manager | Audit Trail | Extra Cost | Best For | |--------|------------|-------------|-----------|---------| | SSE-S3 | AWS fully managed | Difficult | None | General data | | SSE-KMS | Customer via KMS | CloudTrail records | KMS request fee | Compliance-regulated data | | SSE-C | Customer directly | None | None | Highest security requirements |
!S3 server-side encryption options compared
Amazon Redshift Encryption
Redshift encrypts at the cluster level. When you enable encryption at cluster creation, all data stored in the cluster — including all tables, snapshots, and automated backups — is encrypted. Enabling encryption after cluster creation requires a cluster restart, so plan for this upfront.
Using KMS for Redshift encryption gives you key management control while AWS manages the underlying key infrastructure. You define key policies to control which IAM identities can use the key.
CloudHSM (Hardware Security Module) is available for environments that require the highest assurance level. Keys are generated and stored in a dedicated physical hardware device, and they never leave that hardware. Some financial and government regulations explicitly require HSM-backed keys. The cost is significantly higher than KMS, but the hardware guarantee that keys never leave the HSM satisfies the most stringent compliance requirements.
DynamoDB and EMR Encryption
DynamoDB encrypts all stored data by default using AWS-owned keys at no extra cost. If you need more control and the ability to audit key usage, you can switch to a customer-managed CMK. This change can be made without any downtime.
EMR clusters consist of multiple EC2 instances, each with attached EBS volumes. To encrypt these volumes, EMR uses LUKS (Linux Unified Key Setup), which is the standard Linux disk encryption mechanism. You enable LUKS encryption in the EMR security configuration before launching the cluster. Once enabled, all data written to EBS volumes by the cluster nodes is encrypted at the block device level.
Encryption in Transit
Encryption in transit protects data while it moves across a network. Without it, someone monitoring network traffic between your application and AWS could read the data being transferred. Encryption in transit uses TLS (Transport Layer Security) — the same protocol that makes HTTPS secure in your web browser.
For Amazon S3, you enforce HTTPS-only access by adding an aws:SecureTransport condition to the bucket policy. This policy explicitly denies any request that does not use HTTPS, preventing accidental or malicious HTTP access that could expose data in transit.
For Redshift, you set the require_ssl parameter to true in the cluster's parameter group. After applying this change, any connection attempt that does not use SSL is rejected by the cluster. This applies to both BI tool connections and data loading operations.
For EMR, you configure in-transit encryption in the EMR security configuration, covering both the communication between cluster nodes internally and the traffic between the cluster and external services like S3.
Envelope Encryption
Envelope encryption is how KMS actually works under the hood — and understanding it is essential for the DEA-C01 exam. The concept sounds complicated at first, but it solves a simple problem: KMS can only directly encrypt data up to 4 KB in size. Real datasets are gigabytes or terabytes. How do you use KMS for large data?
The solution is to encrypt the data with a data key (which can be any size), and then encrypt the data key with the KMS master key. This is the "envelope" — the data key is the inner envelope, and the KMS key encrypts that inner envelope.
The process step by step:
Step 1: Call KMS GenerateDataKey. KMS returns two things: a plaintext data key and an encrypted data key (the same key encrypted with your KMS master key).
Step 2: Use the plaintext data key to encrypt your actual data locally — on your application or service, not inside KMS. There is no size limit for this step.
Step 3: Store the encrypted data together with the encrypted data key. Immediately discard the plaintext data key from memory. The master key never leaves KMS.
To decrypt: retrieve the encrypted data key and send it to KMS. KMS decrypts it using the master key and returns the plaintext data key. Use that key to decrypt the data locally, then discard the plaintext key again.
The security guarantee: even if an attacker steals the encrypted data and the encrypted data key together, they cannot decrypt anything without access to the KMS master key. And since the master key never leaves KMS, they would need to compromise your AWS account's KMS permissions to get it.
Cross-Account Encryption
In multi-account architectures, a common pattern is one account producing encrypted data and another account consuming it. To allow a different AWS account to use your KMS key, you modify the KMS key policy — not the IAM policy of the consuming account, but the key policy on the key itself. The key policy must explicitly allow the consuming account's IAM role to call kms:Decrypt. Both the key policy permission and the consuming account's IAM role permission must be in place for cross-account decryption to work.
Data Masking and Anonymization
Many data engineering workflows need to share data with teams who should not see the raw sensitive values. You want developers to test with realistic data, analysts to run queries on customer information, but neither group should see actual credit card numbers or social security numbers. This is where masking and anonymization come in.
Masking partially hides the original data while preserving its format. A phone number 010-1234-5678 becomes 010-*-5678. A credit card number 4111-1111-1111-1234 becomes XXXX-XXXX-XXXX-1234. The underlying data in the database is unchanged — masking happens at the display or export layer. Customer service representatives typically see masked card numbers so they can confirm the last four digits with a caller without seeing the full number.
Anonymization irreversibly transforms data so that individuals can no longer be identified. Names, ID numbers, addresses, and other direct identifiers are removed or transformed. Truly anonymized data cannot be linked back to a specific person even with access to external data sources. This is appropriate for research datasets, public data releases, or analytics that only need aggregate patterns.
Pseudonymization replaces identifying information with artificial identifiers (pseudonyms). The real name "John Smith" becomes "USER_4729" in the dataset. Unlike anonymization, pseudonymization is reversible — a separate mapping table links pseudonyms to real identities, and only authorized personnel can access that mapping. GDPR recognizes pseudonymization as a valid data protection measure, though pseudonymized data is still considered personal data under GDPR because re-identification is possible.
AWS tools for implementing masking and anonymization:
AWS Glue DataBrew is a visual data preparation tool that lets you define masking and transformation rules without writing code. You can create a recipe that masks phone number columns, removes name columns, and replaces email addresses with hashed values. When the DataBrew job runs as part of your pipeline, it applies these transformations automatically.
Lake Formation column-level access control functions as a query-time masking mechanism. Rather than physically t