Core Data Concepts

Even if you're brand new to data, this guide walks you through structured, semi-structured, and unstructured data, popular file formats like CSV and Parquet, relational vs. NoSQL databases, and the difference between a data lake and a data warehouse — all explained with everyday analogies.

Azure DP-900 is Microsoft's entry-level data certification. If you're completely new to databases and data storage, this guide is for you. We'll break everything down using everyday analogies so you can understand not just "what" but "why" each concept matters.

 

Data Types: Structured, Semi-Structured, Unstructured

Think of data as different kinds of information you encounter in daily life. Not all information looks the same — some is organized in neat tables, some is loosely structured, and some has no structure at all.

Structured Data

Imagine a hotel guest registry — every guest fills in the exact same form: name, check-in date, room number, phone number. Every row has the same columns. Nothing more, nothing less. This is structured data.

Structured data lives in relational databases and follows a fixed schema (blueprint). It is the most organized type and is easy for computers to search and analyze using SQL.

Examples: customer tables, order records, employee payroll data

Semi-Structured Data

Now imagine a collection of shopping receipts from different stores. Some receipts show a discount, others don't. Some list a loyalty card number, others don't. The format isn't exactly the same every time, but you can still identify things like "item name", "price", and "date" — there is some structure, just not rigidly enforced.

This is semi-structured data. It uses tags or keys to give meaning to values, but allows flexibility in what fields are present. JSON and XML are the most common formats.

Examples: JSON files, XML files, emails with headers

Unstructured Data

Think of a photo album, a collection of voice recordings, or a folder of PDFs. There are no rows or columns. A computer cannot automatically extract "Name: John, Age: 30" from a photo — it just sees pixels. This is unstructured data.

Unstructured data makes up over 80% of all data in the world. Analyzing it requires special techniques like AI and machine learning.

Examples: images, videos, audio files, PDF documents, social media posts

| Type | Analogy | Characteristics | Examples | |------|---------|-----------------|---------| | Structured | Hotel registry | Fixed rows and columns | Customer tables, orders | | Semi-Structured | Shopping receipts | Tags/keys, flexible fields | JSON, XML | | Unstructured | Photo album | No defined structure | Images, videos, PDFs |

 

File Formats: CSV, JSON, Parquet, and More

When data is saved as a file, the choice of format matters a lot. Different formats are built for different purposes. Let's look at each one.

CSV (Comma-Separated Values)

The simplest format of all. Data is stored as plain text, with each value separated by a comma. You can open a CSV file in Notepad or Excel. When you export a spreadsheet from Excel, CSV is often the format it saves to.

Best for: small datasets, sharing data between systems, importing into Excel

JSON (JavaScript Object Notation)

JSON stores data as key-value pairs inside curly braces: . It also supports nested structures — an object inside an object. This makes it perfect for representing complex, flexible data like a user profile with multiple addresses.

When your phone app fetches weather data from the internet, JSON is almost certainly the format being used behind the scenes.

Best for: web APIs, configuration files, semi-structured data

Parquet

Here is where things get interesting for analytics. Parquet stores data column by column, not row by row. Why does this matter?

Imagine you have a table with 1,000,000 rows and 50 columns, and you only want to calculate the average of the "Sales Amount" column. With a row-based format (like CSV), the computer has to read through all 50 columns for every row just to get to the "Sales Amount" values. With Parquet, it jumps directly to the "Sales Amount" column and reads only that. This makes analytics queries dramatically faster and saves storage through efficient compression.

Best for: big data analytics, Azure Data Lake, Synapse Analytics, Apache Spark

ORC (Optimized Row Columnar)

Similar concept to Parquet — columnar storage, great compression. ORC is specifically optimized for the Hadoop and Hive ecosystem, an older but still widely used big data platform.

Best for: Hadoop ecosystem, Hive queries

Avro

Avro stores data row by row and bundles the schema (the description of what fields exist) inside the file itself. The key advantage is schema evolution — if you add a new field to your data, old files that don't have that field can still be read without errors.

Think of Avro like a form letter where the template is attached to every copy — even if you update the template later, older copies are still understandable.

Best for: real-time data streaming (Kafka), event-driven pipelines

XML (eXtensible Markup Language)

XML wraps data in descriptive tags: . It is very readable and self-describing, but the tag overhead makes files large. Many older enterprise systems still use XML for data exchange.

Best for: legacy system integration, configuration files

| Format | Storage Method | Key Advantage | Primary Use | |--------|---------------|--------------|-------------| | CSV | Row-based | Simple, universal | Data exchange | | JSON | Key-value | Flexible structure | APIs, web apps | | Parquet | Column-based | Fast analytics queries | Big data analytics | | ORC | Column-based | Hive-optimized | Hadoop ecosystem | | Avro | Row-based | Schema evolution | Streaming data | | XML | Tag-based | Self-describing | Legacy systems |

 

Database Types: Relational vs Non-Relational

A database is simply an organized place to store and retrieve data. There are two main philosophies.

Relational Databases

Picture a well-organized filing cabinet. Each drawer is labeled (Customer, Orders, Products). Inside each drawer, every document follows the exact same form. Drawers can reference each other — an Order document links to a Customer document via a customer ID. Everything is neat, connected, and consistent.

Relational databases store data in tables (rows and columns) and use SQL to query it. Relationships between tables are defined with Primary Keys (unique identifiers) and Foreign Keys (references to other tables). They are ideal when data consistency and accuracy are critical — banking, healthcare, and inventory systems.

Azure services: Azure SQL Database, Azure Database for MySQL, Azure Database for PostgreSQL

Non-Relational Databases (NoSQL)

Now picture a large, flexible storage bag. You can throw in items of all shapes and sizes. You trade some structure for flexibility and speed at scale. There are four main types of NoSQL databases.

Key-Value Stores: Like a dictionary or a locker room — every item has a unique key (locker number) and you retrieve it instantly with that key. Blazingly fast for simple lookups. Used for caching and session storage. Example: Redis, Azure Cache for Redis

Document Stores: Store data as self-contained documents (like JSON). Each document can have a different structure. Great for content management, user profiles, and product catalogs where each item might have different attributes. Example: MongoDB, Azure Cosmos DB

Column-Family Stores: Group related columns together for efficient storage and retrieval. Designed for massive write throughput — think IoT sensors sending millions of readings per second, or logging systems. Example: Apache Cassandra, Azure Cosmos DB (Cassandra API)

Graph Stores: Store data as nodes (things) and edges (relationships between things). When you need to answer questions like "Who are the friends of friends of Alice?" or "What products do people who bought X also buy?", graph databases excel. Example: Neo4j, Azure Cosmos DB (Gremlin API)

 

Data Lake vs Data Warehouse

When storing massive amounts of data for analytics, you will hear these two terms constantly. They solve different problems.

Data Lake

Picture a natural lake that collects water from rivers, rain, and underground springs — in whatever form it comes, with no filtering. A data lake works the same way. It accepts raw data in any format — structured tables, JSON files, images, video — all poured in together. You only decide how to organize and use the data when you actually need it (Schema-on-Read).

Back to blog list