Data Cataloging Systems

Understand how Glue Data Catalog works like a library index — automatically recording where your data lives and what it looks like, with crawlers and partition sync explained from scratch.

As data volumes grow, the question "where is this data and what does it look like?" comes up constantly. When tens of thousands of files are scattered across S3, opening each one manually is impossible. AWS Glue Data Catalog solves this problem — it is a central metadata store that keeps a map of all your data without actually copying it. DEA-C01 exam questions on this topic ask how you find and manage data at scale.

 

What Is Glue Data Catalog

Think of a library with tens of thousands of books. Without an index card system, finding anything would take forever. With index cards, you instantly know which shelf holds which book. Glue Data Catalog is that index card system for your data.

Glue Data Catalog does not move or copy your actual data files. Instead it records:

Where the data lives (S3 path, RDS endpoint, etc.) What format it is in (CSV, Parquet, JSON, etc.) What columns exist and what type each column holds How the data is partitioned

This information is called metadata. The catalog stores only metadata — your actual data files stay exactly where they are.

 

Databases and Tables Inside the Catalog

The structure inside Glue Data Catalog looks familiar if you have worked with any relational database.

A database is a logical grouping of related tables. For example, a database called "sales_db" might contain tables named "orders", "customers", and "products".

A table defines the schema (structure) and location of actual data. Each table definition includes:

| Field | Example | Purpose | |-------|---------|----------| | Location | s3://my-bucket/orders/ | Where the data files are | | Format | Parquet | File type | | Columns | order_id (int), amount (double) | Column names and data types | | Partition keys | year, month | How data is split |

These are virtual tables — they do not store any rows. They are definition documents that tell query engines where to look and how to read the data.

 

Integration with Athena, Redshift Spectrum, and EMR

The biggest strength of Glue Data Catalog is that multiple services share the same metadata. Register a table once, and every service can use it without any extra configuration.

Imagine you have order data stored in S3. You want to analyze it:

Athena: Run SQL queries directly against S3 — serverless, no extra setup Redshift Spectrum: Query S3 data as an external table from inside Redshift EMR: Process the data at massive scale using Spark or Hive

All three services look at the same table definition in Glue Data Catalog. There is no need to register the same dataset three times in three different places.

 

Glue Crawlers — Automated Schema Discovery

Manually entering metadata into the catalog is tedious, especially when files are added daily or schemas change frequently. Glue Crawlers automate this work.

A crawler is like an automated detective. You point it at a data source — an S3 bucket, an RDS database, a DynamoDB table — and it goes in, figures out the structure, and registers everything in the catalog.

Key crawler capabilities:

Scheduled runs: Schedule crawlers to run hourly, daily, or weekly. New data and schema changes are detected automatically. Classifiers: The crawler looks at each file and automatically determines its format — CSV, JSON, Parquet, ORC, and more. You do not need to tell it. For unusual formats, you can write a custom classifier. Partition detection: If your S3 folders follow a pattern like year=2026/month=03/day=15, the crawler recognizes this as partitioning and registers all partitions in the catalog.

The crawler workflow step by step:

Scan the specified S3 path (or other source) Identify file formats and structure using classifiers Compare with existing catalog tables Create new tables or update existing table schemas with detected changes Add any newly discovered partitions to the catalog

 

Partition Sync — Three Methods

Partitioning means storing data split into folders by a specific attribute — for example, storing log data in daily folders so you can read only the data you need for a specific date range.

The problem: when a new partition folder appears in S3, the catalog does not automatically know about it. Someone or something has to tell the catalog that a new folder exists. This is partition synchronization.

Method 1 — Re-run the Crawler: The simplest approach. Run the crawler again and it detects new partitions, then adds them to the catalog. The downside is that the crawler re-scans the entire source, which takes time. Best when new partitions are added infrequently.

Method 2 — MSCK REPAIR TABLE: An SQL command you run from Athena:

This works when your S3 folders follow Hive-compatible naming (key=value format). One command syncs all new partitions into the catalog. Faster than re-running a crawler, but only works with Hive-format folder structures.

Method 3 — BatchCreatePartition API: Directly call the Glue API in code, specifying exactly which partitions to add. This is the fastest method and handles thousands of new partitions at once. Ideal for high-frequency environments where new partitions appear every hour. Easily automated inside a Lambda function or a Glue job.

| Method | Speed | Best For | |--------|-------|----------| | Crawler re-run | Slow | Infrequent partition additions | | MSCK REPAIR TABLE | Medium | Hive-format, manual sync | | BatchCreatePartition API | Fast | High-frequency large-scale partitioning |

 

Replacing Hive Metastore with Glue Catalog

EMR (Elastic MapReduce) is AWS's managed big-data service based on open-source Hadoop and Spark. Traditionally, EMR clusters used their own built-in Hive Metastore to track schema information. The problem: when you terminate the cluster, that metastore disappears with it.

By replacing the Hive Metastore with Glue Data Catalog, you gain:

Metadata persists even after the cluster is shut down EMR, Athena, and Redshift Spectrum all share the same metadata No need to register tables separately in each service

To enable this, simply turn on the "Use Glue Data Catalog as the metastore" option when creating an EMR cluster. After that, any table you create in Hive or Spark SQL is automatically registered in the Glue catalog.

 

Exam Key Points

"Central metadata store" → Glue Data Catalog "Auto-discover and register data schemas" → Glue Crawlers "Add many new S3 partitions quickly in code" → BatchCreatePartition API "Manually sync partitions from Athena" → MSCK REPAIR TABLE "EMR and Athena share the same metadata" → Glue Data Catalog replacing Hive Metastore "Automatically detect file formats" → Crawler Classifiers The catalog stores only metadata, not the actual data files

Back to blog list