The DP-900 exam often asks which Cosmos DB API to use in a given scenario and which Blob Storage tier is most cost-effective. This guide explains why Azure has so many different storage services and when to use each one — with simple analogies.
Why Do We Need Different Storage Services?
Think about different kinds of storage in real life. A convenience store keeps its best-selling items at eye level, easy to grab. A warehouse stores bulk goods less accessibly but more cheaply. A deep-archive facility keeps rarely-accessed records in cold storage at the lowest cost, but retrieving them takes time.
Cloud storage works the same way. Different types of data — files, logs, analytical datasets, globally distributed app data — need different storage solutions optimized for their access patterns and costs.
Azure's main storage services at a glance:
| Service | Analogy | Best For | |---------|---------|---------| | Blob Storage | Large file locker | Photos, videos, documents, backups | | Azure Files | Shared network drive | Multi-user file sharing | | Table Storage | Simple cloud spreadsheet | Key-value structured data | | Data Lake Gen2 | Massive analytics warehouse | Raw data for analysis | | Cosmos DB | Global library network | Globally distributed NoSQL data |
Azure Blob Storage — The File Locker
Blob stands for Binary Large Object. It is Azure's object storage for unstructured data: images, videos, PDFs, backups — anything that does not fit neatly into rows and columns. Think of Blob Storage as a giant locker room where you can put anything as long as it fits.
Three Blob Types
Block Blob is the most common type. Files are broken into blocks that can be uploaded in parallel, making it fast for large files. Use it for images, videos, documents, and backups.
Append Blob only allows data to be added at the end — you cannot modify existing data. Perfect for log files where you continuously append new log entries.
Page Blob stores data in 512-byte pages optimized for random read/write. Used by Azure VM disks (VHD files).
!3 blob storage types
Access Tiers — Balancing Cost and Speed
Imagine a filing cabinet. The top drawer holds files you use every day. The middle drawer holds files you need occasionally. The bottom drawer holds files you rarely open. The deeper the drawer, the cheaper to store, but the longer it takes to retrieve.
| Tier | Storage Cost | Access Cost | Best For | |------|-------------|-------------|---------| | Hot | High | Low | Data accessed daily | | Cool | Medium | Medium | Data kept 30+ days, accessed occasionally | | Cold | Low | Medium-High | Data kept 90+ days, accessed rarely | | Archive | Very Low | Very High + delay | Data kept 180+ days, almost never accessed |
Archive tier important note: to read archived data, you must first rehydrate it to the Hot or Cool tier. This can take up to 15 hours. It is the cheapest storage but has no instant access.
Redundancy Options
LRS (Locally Redundant Storage) keeps three copies within the same datacenter — lowest cost. ZRS (Zone-Redundant Storage) spreads copies across three availability zones in the same region. GRS (Geo-Redundant Storage) replicates asynchronously to a second region for disaster recovery. GZRS combines zone redundancy with geo-redundancy for the highest durability.
Azure Files — Shared Network Drive
Azure Files provides fully managed file shares in the cloud that you can mount just like a traditional file server on your PC. It supports SMB (Windows file sharing protocol) and NFS, so Windows, Linux, and macOS clients can all connect.
The biggest use case is lifting and shifting on-premises file servers to the cloud, or sharing configuration files and scripts across multiple virtual machines.
Azure Table Storage — Simple Key-Value Store
Azure Table Storage is a NoSQL key-value store. Think of it as a giant spreadsheet in the cloud where each row can have different columns — there is no fixed schema. It is much cheaper and more scalable than a relational database, but it does not support SQL JOINs or complex queries.
Note: Azure Table Storage and Cosmos DB Table API look similar but differ. The Cosmos DB Table API offers global distribution, single-digit millisecond latency, and a stronger SLA.
Azure Data Lake Storage Gen2 — Analytics Warehouse
Data Lake Storage Gen2 (ADLS Gen2) is built on top of Blob Storage with a hierarchical file system added. It is designed to store massive amounts of raw data — from terabytes to petabytes — and let analytics engines like Azure Databricks, Synapse Analytics, and HDInsight read it directly and efficiently.
If Blob Storage is a large warehouse, ADLS Gen2 is a meticulously organized library with folders, subfolders, and fine-grained access control (POSIX-compatible ACLs) at the file and folder level.
Azure Cosmos DB — Global Library Network
Why Global Distribution Matters
Imagine your database server sits in Seoul. A user in New York makes a request, and the data has to travel across the Pacific Ocean and back — that round trip alone can take 200ms or more. Now imagine if every major city had its own library branch. The New York user checks out the closest branch and gets an instant response.
Cosmos DB automatically replicates your data across Azure regions worldwide, guaranteeing single-digit millisecond latency no matter where your users are. It also supports multi-master writes, meaning users in Seoul and New York can write data simultaneously.
Core Concepts
Request Unit (RU): the unit of throughput in Cosmos DB. Reading a 1 KB item costs 1 RU. Writes cost more. You provision how many RUs per second your database needs.
Partition Key: the attribute used to distribute data across partitions. Like organizing books by genre so the librarian can find them instantly. Choose a partition key with high cardinality (many unique values) so data spreads evenly and you avoid hot partitions.
Consistency Levels: how strictly Cosmos DB synchronizes data across global replicas. Strong guarantees always-current data but adds latency. Eventual is fastest but may briefly return slightly stale data. Session, Bounded Staleness, and Consistent Prefix offer middle-ground options.
Cosmos DB APIs
Cosmos DB runs one engine but offers multiple APIs so you can use familiar database drivers without rewriting your application.
| API | Data Model | Use Case | |-----|-----------|---------| | NoSQL (Core) | JSON documents | General purpose, the default choice | | MongoDB | BSON documents | Migrating MongoDB apps | | Cassandra | Column families | High-volume writes, time-series data | | Gremlin | Graph (nodes + edges) | Social networks, recommendation engines | | Table | Key-value | Migrating Azure Table Storage apps |
The NoSQL API is the native Cosmos DB API and the recommended default. The others exist for compatibility when migrating from existing databases.
Exam Key Points
"images, videos, file storage" -- Blob Storage "daily access, higher storage cost" -- Hot tier "kept 30+ days, occasional access" -- Cool tier "almost never accessed, rehydration up to 15 hours" -- Archive tier "three copies in same datacenter" -- LRS / "copies in another region" -- GRS "log files, append-only" -- Append Blob "VM disk (VHD)" -- Page Blob "shared file server, SMB/NFS protocol" -- Azure Files "simple key-value, no fixed schema" -- Azure Table Storage "massive analytics data, hierarchical namespace" -- Data Lake Storage Gen2 "global distribution, single-digit millisecond latency" -- Cosmos DB "JSON documents, default API" -- Cosmos DB NoSQL API "migrating MongoDB apps" -- Cosmos DB MongoDB API "graph data, social networks" -- Cosmos DB Gremlin API "column families, high-volume writes" -- Cosmos DB Cassandra API "data distribution key, even spread" -- Partition Key "throughput unit, 1KB read = 1RU" -- Request Unit (RU)