Azure Cosmos DB is a globally distributed NoSQL database that lets you read and write data quickly from anywhere in the world. For the AZ-204 exam, the key topics are SDK data operations, partition key design, 5 consistency levels, change feed, and server-side programming.
SDK: Accessing Data Through a Hierarchy
When you first encounter the Cosmos DB SDK, the many class names can be confusing. Think of a company's organizational chart to make it easier. Under the headquarters (CosmosClient) there are departments (Database), under departments there are teams (Container), and each individual team member is an item (Item).
| Class | Role | Analogy | |-------|------|---------| | CosmosClient | Connects to the entire Cosmos DB account | Main switchboard | | DatabaseProxy | Reference to a specific database | Specific department | | ContainerProxy | Reference to a specific container | Specific team | | ItemProxy | Individual document (JSON data) | Individual team member |
With the SDK, you use a single connection string and navigate down this hierarchy to work with data.
Basic CRUD Operation Pattern
Connect: CosmosClient(endpoint, credential) Database reference: client.get_database_client("mydb") Container reference: db.get_container_client("mycontainer") Create item: container.create_item(body={"id": "1", "pk": "user1", ...}) Read item: container.read_item(item="1", partition_key="user1") Update item: container.upsert_item(body={...}) Delete item: container.delete_item(item="1", partition_key="user1") Query: container.query_items(query="SELECT * FROM c WHERE c.type = 'order'", ...)
Every read/write operation consumes RU (Request Unit) cost. Reading a 1KB item costs about 1 RU; writes cost more RUs.
Partition Key: The Criterion for Dividing Data
The partition key is the basis Cosmos DB uses to distribute data across multiple servers. Think of a library. If you organize books alphabetically, books starting with certain letters pile up on one shelf. If you organize by genre instead, they spread evenly across shelves. Cosmos DB works the same way.
You must remember the conditions for a good partition key.
| Condition | Description | Examples | |-----------|-------------|---------| | High cardinality | The more unique values the better | User ID, Order ID | | Uniform distribution | Data should not cluster around specific values | City, Category | | Matches queries | Should match fields you frequently filter on | Fields used in WHERE clauses |
You also need to know examples of bad partition keys.
Date: Data clusters on specific dates (hot partition problem) Boolean: Only true/false means only 2 partitions Fixed categories: Too few distinct values concentrate load on specific partitions
Partition keys cannot be changed after item creation, so choose carefully during initial design.
5 Consistency Levels: Balancing Speed and Accuracy
In a distributed database, data is replicated across servers worldwide. If you write data in Seoul, the same data needs to reach servers in Tokyo or New York. The problem is this transfer takes time. Consistency levels control the balance between "how fresh is the data guaranteed to be" and "how fast does it respond."
Think of breaking news as an example. When an event occurs, some outlets report after verification for accuracy, while others report quickly but correct later.
| Consistency Level | Guarantee | Performance | Use Case | |-------------------|-----------|-------------|---------| | Strong | Always guarantees latest data. Blocks reads until all replicas reflect the write | Slowest | Financial transactions, inventory | | Bounded Staleness | Guarantees data within K versions or T seconds | Slow | Leaderboards, reservation systems | | Session | Within the same session, always reads your own writes (default) | Medium | Shopping cart, user profile | | Consistent Prefix | Guarantees order but not freshness | Fast | Social feeds, comments | | Eventual | Will eventually converge, no immediate guarantee | Fastest | Like counts, view counts |
!Cosmos DB's 5 consistency levels
The most frequently tested topics are that Session is the default and the characteristics of each level.
Change Feed: Real-Time Event Processing
Change Feed is a feature that streams changes happening in a Cosmos DB container in real time. Think of convenience store CCTV. The moment someone picks up a product or checks out is recorded in real time, and other systems (inventory management, security) can process this immediately.
Key characteristics of change feed you must remember.
Supported operations: Only detects inserts (INSERT) and updates (UPDATE) Unsupported operations: Deletes (DELETE) are not detected by default (use TTL instead) Processing methods: Azure Functions trigger, Change Feed Processor library Use cases: Real-time notifications, event sourcing, cache invalidation, analytics pipelines
Integration with Azure Functions
Using a Cosmos DB trigger, a function automatically executes whenever a new item is added or an existing item is changed in a container. For example, you can set up a workflow that automatically decrements inventory the moment an order comes in, without writing additional coordination code.
Server-Side Programming: Logic Executing Inside the Database
Cosmos DB supports three mechanisms for executing code directly inside the database server. Like a factory assembly line, processing data where it lives eliminates the need to send it back and forth over the network.
Stored Procedure
Bundles multiple operations into a single transaction. Using bank transfer as an example, debiting account A and crediting account B must either both succeed or both fail together. Stored procedures guarantee this atomic behavior.
Written in JavaScript Transaction guarantee applies only within the same partition key Full rollback on failure
Trigger
Code that automatically executes before or after item creation, modification, or deletion.
Pre-trigger: Executes before an item is saved (validation, data normalization) Post-trigger: Executes after an item is saved (logging, sending notifications)
UDF (User-Defined Function)
A custom function you can call inside SQL queries. For example, you can write a tax calculation logic as a UDF and call it directly from queries. Does not support transactions and is read-only.
| Type | Transaction | Execution Timing | Purpose | |------|-------------|-----------------|---------| | Stored Procedure | Supported | Explicit call | Atomic multi-operation | | Pre-trigger | Not supported | Before write operation | Validation, normalization | | Post-trigger | Not supported | After write operation | Logging, notifications | | UDF | Not supported | Called within query | Custom calculation logic |
Exam Key Points
"CosmosClient -> Database -> Container -> Item" -- SDK hierarchy (memorize the order)
"Using date or boolean as partition key" -- Bad examples (causes hot partition)
"High cardinality + uniform distribution" -- Good partition key conditions
"Default consistency level" -- Session
"Always guarantees latest data, slowest" -- Strong
"Eventually consistent, fastest" -- Eventual
"What change feed detects" -- Only inserts and updates (deletes not supported by default)
"Atomic multi-operation, transaction guarantee" -- Stored procedure
"Validation before write" -- Pre-trigger
"Custom calculation inside queries" -- UDF (no transaction support)