Data has a lifespan. Log files created today will be queried frequently tomorrow, but logs from three years ago are rarely opened. Storing both in the same high-cost storage tier is wasteful. Data lifecycle management is about planning how long to keep data and in what form. Schema evolution is the set of techniques that let you change data structures without breaking existing data. DEA-C01 exam questions on these topics focus on cost optimization and operational stability.
S3 Lifecycle Policies
S3 offers multiple storage classes. Moving data to cheaper classes as access frequency declines can dramatically reduce storage costs.
| Storage Class | Characteristics | Best For | |---------------|----------------|----------| | S3 Standard | Fast access, highest cost | Frequently accessed recent data | | S3 Standard-IA | Occasional access, medium cost | Data after 30 days | | S3 Glacier Instant Retrieval | Rare access, low cost, instant retrieval | Data after 90 days | | S3 Glacier Flexible Retrieval | Very rare access, very low cost, minutes-to-hours retrieval | Long-term archiving | | S3 Glacier Deep Archive | Almost never accessed, lowest cost, 12-hour retrieval | Compliance data held for years |
Lifecycle policies automate these transitions. For example, you could configure:
After 30 days: Standard → Standard-IA automatically After 90 days: Standard-IA → Glacier Instant Retrieval After 365 days: Glacier Deep Archive After 2,555 days (7 years): Auto-delete
Transition moves data to a cheaper class. Expiration deletes data automatically. Lifecycle policies can apply to an entire bucket, a specific prefix (folder path), or objects with specific tags.
DynamoDB TTL — Automatic Expiry of Records
DynamoDB is a NoSQL database commonly used for session data, temporary tokens, shopping carts, and other data that becomes irrelevant after a certain time.
TTL (Time To Live) lets you set an expiration time on individual items. When an item's expiration time passes, DynamoDB automatically deletes it — no deletion code required on your part.
How it works: Designate an attribute on your table to hold the expiration time (for example, ) Store a Unix timestamp in that attribute for each item DynamoDB scans for expired items in the background and deletes them Deletion typically happens within 48 hours after expiration
Important notes: TTL deletion is free (no write capacity consumed) Items are not deleted at exactly the expiration moment, so your application should check the expiration field itself Combine DynamoDB Streams with TTL to capture delete events for additional processing
Redshift Data Management
Redshift is a petabyte-scale data warehouse. Efficiently loading data in and moving data out is critical.
The COPY command loads data in bulk from S3, DynamoDB, EMR, and other sources into Redshift. It is far faster than individual INSERT statements.
The UNLOAD command exports data from Redshift to S3. Use it to save analysis results or hand data off to another system.
Snapshots are full backups of a Redshift cluster. Automatic snapshots run every 8 hours or every 5 GB of changes, whichever comes first. Manual snapshots run on demand. Use snapshots to restore a cluster in another region for disaster recovery.
Schema Design Principles
How you arrange data has a huge impact on query performance.
Redshift DISTKEY and SORTKEY:
Redshift distributes data across multiple nodes. DISTKEY specifies which column to use as the distribution key. When two tables are joined on a column that is the DISTKEY for both, the matching rows are already on the same node — no network transfer required, which speeds up joins significantly.
SORTKEY specifies the order in which data is stored within each block on disk. If your queries frequently filter by date range, making the date column the SORTKEY lets Redshift skip entire blocks outside the target range — dramatically speeding up range queries.
DynamoDB Partition Key: DynamoDB distributes data across partitions based on the partition key. If too much data lands on one partition key value, that partition becomes a hot partition, causing performance bottlenecks. Choose a partition key with high cardinality (many unique values). User ID is a much better partition key than gender.
S3 Partitioning: Decide how to organize your S3 folders when writing data. When Athena queries partitioned data, it skips irrelevant partitions entirely, reducing the data scanned — which lowers both cost and query time.
Example:
Schema Evolution — When Your Data Structure Changes
Businesses change. The data structure you design today may not fit six months from now. Schema evolution is the ability to change a schema without breaking existing stored data.
Apache Iceberg: A table format layer that sits on top of data files in S3. It supports adding columns, changing column types, and dropping columns. It also provides time-travel queries, letting you query data as it existed at a past point in time. Natively supported by AWS Glue and Athena.
Apache Avro: A serialization format that stores the schema alongside the data. Even when the schema changes, data written with the old schema can still be read correctly. Frequently combined with Kafka for schema evolution in streaming pipelines.
SCT (Schema Conversion Tool) and DMS (Database Migration Service): Used for migrations between different database engines. SCT converts schemas from Oracle, SQL Server, and others to formats compatible with Redshift or Aurora. DMS then migrates the actual data from the old database to the new one.
Data Lineage — Tracking Your Data's Journey
Data lineage tracks where data came from, what transformations it went through, and where it ended up. It is essential for regulatory compliance, auditing, and debugging.
Imagine a report shows a wrong number. Without lineage, you must check each step: was it the source data? The ETL transformation? The aggregation query? With lineage, you can trace the data's path backwards and find the problem quickly.
AWS services that help implement lineage: AWS Glue: Automatically records job execution history and transformation steps Amazon DataZone: Visualizes lineage across the organization's entire data catalog CloudTrail: Audit log at the API call level Lake Formation: Access records across the entire data lake
Exam Key Points
"Automatically move aging data to cheaper storage" → S3 Lifecycle policies "Auto-delete expired DynamoDB items" → TTL "Bulk load data from S3 to Redshift" → COPY command "Export data from Redshift to S3" → UNLOAD command "Optimize Redshift join performance" → DISTKEY "Optimize Redshift range queries" → SORTKEY "Read old data after schema changes" → Apache Iceberg or Avro "Migrate on-premises DB to AWS" → SCT + DMS "Track data movement and transformation history" → Data Lineage