Why Data Ingestion Is the Foundation of Every ML Project
If you've spent any real time building production ML systems, you know the uncomfortable truth: most of your time is spent on data, not models. Industry surveys consistently show that data preparation consumes 70-80% of a data scientist's time. The model architecture matters, but it's a distant second to having clean, relevant, and well-structured data flowing into your training pipeline.
AWS MLA-C01 takes this reality seriously. Domain 1 — Data Preparation — is the most foundational section of the exam, and data ingestion and storage is its cornerstone. This post walks through the AWS services and architectural patterns you need to know, with the practical context that makes them stick.
Amazon S3 as Your ML Data Lake Hub
S3 is not just a file storage service — it is the connective tissue of the entire AWS ML ecosystem. SageMaker reads training data from S3. Glue ETL writes transformed data to S3. Athena queries S3 directly. Kinesis Firehose delivers streaming data to S3. Every major AWS ML service integrates with S3 at its core.
For ML workloads, how you organize and configure S3 matters as much as using it. Partitioning your data with a logical prefix structure (for example, ) allows query engines like Athena and Glue to skip irrelevant partitions entirely — a technique called partition pruning that can reduce query costs by orders of magnitude on large datasets.
For training datasets, the standard convention is to separate data into , , and prefixes. SageMaker Training Jobs can point directly to each prefix as a separate data channel, making the data pipeline clean and reproducible.
Storage class selection is a cost optimization lever. Active training data and recent experiment outputs belong in S3 Standard. Older model artifacts and infrequently accessed datasets suit S3 Intelligent-Tiering or S3 Standard-IA. Long-term compliance archives belong in S3 Glacier. S3 Lifecycle policies automate these transitions, so you never have to manage them manually.
S3 Select is a lesser-known but powerful feature. It allows server-side filtering of CSV, JSON, and Parquet files using SQL syntax, so you only transfer the rows and columns you actually need. Combined with SageMaker Processing, it can dramatically reduce preprocessing costs for large datasets.
Data Formats for ML Workloads
Choosing the right data format affects training speed, storage cost, and compatibility with AWS services. Here is a practical breakdown:
| Format | Key Characteristic | Best For | |--------|-------------------|----------| | CSV | Human-readable, universal | Prototyping, small datasets | | Parquet | Columnar, compressed, schema-embedded | Large analytical datasets, Athena/Glue | | RecordIO | SageMaker native, streaming-friendly | Built-in algorithms (XGBoost, image classifiers) | | TFRecord | TensorFlow native | TF/Keras training pipelines | | ORC | Optimized for Hive/EMR | Spark-based large-scale batch | | JSON Lines | Flexible, handles nested schemas | NLP, event logs, semi-structured data |
The most common pattern in production is to ingest raw data in CSV or JSON, transform it to Parquet using Glue ETL or EMR, and store the Parquet files in the S3 data lake. When feeding SageMaker built-in algorithms, check the algorithm documentation — many accept both CSV and RecordIO, but some require a specific format.
Batch Ingestion Patterns
Batch ingestion handles large volumes of data on a scheduled or triggered basis. AWS offers several services depending on the scale and complexity of your pipeline.
AWS Glue is the go-to service for serverless ETL in the AWS ML stack. A Glue Crawler automatically discovers schema from S3, RDS, Redshift, and DynamoDB, registering the metadata in the Glue Data Catalog. Glue ETL Jobs run PySpark-based transformations at scale. The Visual ETL interface lets you build transformation pipelines without writing code, which is useful for common operations like joins, deduplication, and type casting. Glue is also deeply integrated with SageMaker, making it natural to use as the data prep layer before a training job.
Amazon EMR is for workloads that outgrow serverless ETL. When you need fine-grained Spark configuration, specialized Hadoop ecosystem tools, or multi-step distributed processing across terabytes of data, EMR gives you full control. EMR Serverless removes cluster management while preserving Spark's power — a good middle ground.
AWS DMS handles database replication and migration. It is the right tool for continuously syncing an operational database to S3 for ML purposes, using Change Data Capture to capture new and updated rows in near-real-time.
Streaming Ingestion Patterns
When your ML system needs to react to data as it arrives — fraud detection, real-time recommendations, anomaly detection — you need a streaming ingestion architecture.
Amazon Kinesis Data Streams provides durable, scalable streaming with shard-based throughput control. Each shard handles up to 1 MB/second of input or 1,000 records/second. Data is retained for up to 365 days, enabling replay. Consumers include Kinesis Data Analytics (for real-time SQL or Flink processing), Lambda, and custom KCL applications.
Amazon Kinesis Data Firehose is the simplest path from a data stream to a persistent destination. You configure a source (Kinesis streams, direct PUT, MSK), a destination (S3, Redshift, OpenSearch, Splunk), and optionally a Lambda function for in-flight transformation. Firehose handles buffering, compression, and batching automatically. In ML pipelines, it is commonly used to land event data in S3 for later batch training or to feed a feature store.
Amazon MSK (Managed Streaming for Apache Kafka) is the right choice when your team has existing Kafka expertise or tooling. It gives you full Kafka semantics with AWS managing the brokers, ZooKeeper, and storage. MSK Connect supports Kafka connectors for streaming data to and from dozens of systems.
!Batch versus streaming ingestion
AWS Lake Formation — Governance at Scale
As data lake usage grows across teams, managing access through IAM bucket policies becomes unwieldy. Lake Formation provides fine-grained access control at the database, table, column, and row level, layered on top of the Glue Data Catalog. A data scientist can be granted SELECT on specific columns of a specific table without needing to understand the underlying S3 bucket structure. Sensitive columns like PII can be automatically masked for certain roles. For organizations building shared ML data lakes, Lake Formation is the right governance layer.
SageMaker Training Input Modes
This is a concrete technical topic that appears regularly in MLA-C01 questions. SageMaker supports three input modes for reading training data from S3.
File Mode copies all data from S3 to the local storage of the training container before training begins. It is the simplest mode and works well when your dataset fits comfortably in local storage. The downside is initialization latency for large datasets.
Pipe Mode streams data from S3 to the training container through a Unix named pipe. There is no local copy — data flows directly from S3 as the training loop requests it. This eliminates disk storage costs and startup latency. It works best with RecordIO format and is ideal for large datasets that would be expensive to copy.
FastFile Mode mounts S3 data as a high-performance file system using Amazon FSx for Lustre under the hood. It behaves like File Mode (standard file system access patterns), but data is fetched lazily as accessed rather than copied upfront. It supports random access patterns, making it more flexible than Pipe Mode, and is the generally recommended mode for new workloads.
Exam Tips
"Streaming data, auto-delivery to S3 with optional transformation" -- Kinesis Data Firehose with Lambda "Manual shard management, custom stream processing logic" -- Kinesis Data Streams "Serverless ETL, automatic schema discovery" -- AWS Glue with Glue Crawler "Large-scale distributed Spark, full cluster control" -- Amazon EMR "Serverless SQL on S3" -- Amazon Athena "Column/row-level access control on data lake" -- AWS Lake Formation "All training data pre-copied before training starts" -- File Mode "No disk copy, streaming read during training, RecordIO preferred" -- Pipe Mode "Lazy file system mount, supports random access" -- FastFile Mode "Columnar format, best for analytical queries" -- Parquet "SageMaker built-in algorithm default format" -- RecordIO or CSV (check per algorithm) "Database to AWS migration with CDC" -- AWS DMS "Accelerate large uploads across geographies" -- S3 Transfer Acceleration