DynamoDB Data Modeling and Queries

Design DynamoDB partition keys, GSI/LSI, Query vs Scan, consistency models, and transactions.

DynamoDB is the second most tested service on the AWS DVA-C02 exam. Understanding key design and query patterns is essential for passing.

 

What Exactly Is DynamoDB?

Traditional databases like RDS or MySQL store data in fixed table structures with rows and columns. DynamoDB is a NoSQL database — it is flexible in structure and blazingly fast (millisecond response times).

Think of it this way: if a traditional database is like a spreadsheet, DynamoDB is like a bundle of sticky notes. Each sticky note (item) can have different content, and you can find any one of billions of notes in an instant.

 

Primary Key Design — The Foundation of How Data Gets Stored

The most important concept in DynamoDB is the Primary Key. The primary key determines which server (partition) will store your data.

Partition Key

The partition key decides how data is distributed across multiple servers. Think of it like postal codes at a post office — the more evenly postal codes are distributed, the faster deliveries happen.

The key principle: always choose an attribute with many unique values (high cardinality) as your partition key.

Bad example: using a date as the partition key. Thousands of records created on the same day all land on the same server, causing a traffic jam. This is called a hot partition problem.

Good example: using user ID as the partition key. Each user's data is spread across different servers, distributing the load evenly.

Sort Key

The sort key orders data within the same partition. Using both a partition key and sort key together enables range queries.

For example, with user ID as the partition key and timestamp as the sort key, you can quickly retrieve a specific user's recent activity in chronological order.

 

Indexes — When You Need to Search in Multiple Ways

With only the primary key, you can only search data one way. When you want to search quickly by other attributes, you use indexes. It works just like a book index — it helps you find what you need without reading the entire book.

GSI (Global Secondary Index)

A GSI lets you query using completely different partition and sort keys than the original table. You can add a GSI at any time after the table is created, which makes it very flexible. However, it requires separate provisioned capacity.

Example: if your table is organized by user ID but you also want to search by email address, create a GSI with email as the partition key.

LSI (Local Secondary Index)

An LSI keeps the same partition key as the original table but uses a different sort key. The critical limitation: an LSI can only be added when you first create the table. You cannot add one later, so plan ahead.

 

Query vs Scan — The Key to Efficient Data Retrieval

Query

Use Query when you know the partition key. It searches only the relevant partition, making it highly efficient and fast. Think of it as walking directly to a specific section in a library to find a book.

Scan

Scan reads through the entire table from start to finish. It can find any data, but since it reads everything, it is slow and expensive. Think of it as opening every single book in the library to find what you need.

Avoid Scan whenever possible. If Scan is unavoidable, use ProjectionExpression to fetch only the attributes you need, and consider Parallel Scan to speed things up by processing multiple sections at once.

!Query versus Scan in DynamoDB

Consistency Models — How Fresh Do You Need Your Data?

After writing data to DynamoDB, there can be a brief moment before all servers reflect the update. DynamoDB offers two read options to handle this.

Eventually Consistent (default): DynamoDB's standard mode. There may be a slight delay before the latest data is visible. This costs half as much and is sufficient for most use cases.

Strongly Consistent: guarantees you read the absolute latest data immediately. Use the ConsistentRead=True option to enable it. It costs twice as much, but is necessary for scenarios like financial transactions where stale data is not acceptable.

 

Transactions — When Multiple Operations Must Succeed or Fail Together

Imagine a bank transfer: money must be deducted from Account A and added to Account B at the same time. Both operations must either both succeed or both fail. This is exactly what transactions are for.

TransactWriteItems: processes write operations across multiple items atomically — all succeed or all fail.

TransactGetItems: reads multiple items atomically.

Transactions cost twice as much as regular operations. Use them only when atomicity is truly required.

 

Batch Operations — Processing Multiple Items at Once

BatchWriteItem: write or delete up to 25 items in a single request.

BatchGetItem: read up to 100 items in a single request.

The key difference from transactions: batch operations are not atomic. Some items may succeed while others fail. Failed items are returned in UnprocessedItems and must be retried.

 

Exam Key Points

"Efficient query key design" -- High cardinality partition key

"Query with different attributes, can be added after table creation" -- GSI

"Same partition key, different sort key, added at creation only" -- LSI

"Efficient data retrieval" -- Query (partition key-based)

"Reads entire table (inefficient)" -- Scan

"Need immediate latest data" -- Strong consistency (ConsistentRead=True)

"Atomic writes across multiple items" -- TransactWriteItems

"Bulk write, partial failures possible" -- BatchWriteItem (check UnprocessedItems)

To avoid hot partitions, choose a partition key with high cardinality

Back to blog list