Data ingestion is the very first step in any data engineering pipeline. Just as a factory needs raw materials before it can make products, a data system needs to collect data before it can analyze anything. In the AWS DEA-C01 exam, a major focus is "which ingestion service do you choose for a given scenario?" This guide starts from the two core concepts — streaming and batch — and explains each AWS service with beginner-friendly analogies.
Streaming vs Batch — A Water Supply Analogy
There are two fundamental ways to collect data.
Streaming is like a water pipe: data flows continuously in real time. Every time a user clicks in an app, every time a sensor reads a temperature, that moment's data is sent immediately. Latency is in the milliseconds-to-seconds range, and it answers the question "what is happening right now?"
Batch is like a water tank: data accumulates and is delivered all at once on a schedule. Think of processing a full day's transactions at midnight, or analyzing last week's sales every Monday morning. There is no real-time aspect, but it is efficient and simpler to implement.
| Aspect | Streaming | Batch | |--------|-----------|-------| | Data arrival | Real-time (continuous) | Periodic (all at once) | | Latency | Milliseconds to seconds | Minutes to hours to days | | Use cases | Real-time fraud detection, stock price monitoring | Monthly sales reports, nightly ETL jobs | | AWS services | Kinesis Data Streams, MSK | DMS, AppFlow, Transfer Family |
!Streaming versus batch data ingestion
Kinesis Data Streams — The Highway for Real-Time Data
Kinesis Data Streams is a service for collecting and processing large volumes of real-time data. Picture a highway with multiple lanes (shards). The more lanes you have, the more cars (data records) can travel at the same time.
Understanding shards is the key concept here. Each shard provides 1 MB/s write throughput and 2 MB/s read throughput. As data volume grows, you split shards to increase capacity.
The partition key determines which shard a record goes to. Records with the same partition key always go to the same shard. For example, if you use a customer ID as the partition key, all events from the same customer land in the same shard in order.
Key features:
Retention: By default data is kept for 24 hours; this can be extended up to 365 days. Even if a downstream processor fails, the data is still there waiting. Replayability: You can read the same data multiple times. This means you can reprocess data after fixing a bug in your processing logic, or have multiple independent systems reading the same stream simultaneously. Multiple consumers: AWS Lambda, Kinesis Data Firehose, and KCL (Kinesis Client Library) applications can all connect as consumers at the same time. Enhanced Fan-out: In standard mode, all consumers share the 2 MB/s read throughput per shard. Enhanced Fan-out gives each registered consumer its own dedicated 2 MB/s, so multiple high-throughput consumers do not interfere with each other.
Kinesis Data Streams requires hands-on management. You configure the number of shards yourself and build your own consumer applications. In return you get complete flexibility and full replayability.
Kinesis Data Firehose — Automatic Delivery to Your Destination
Kinesis Data Firehose is like a courier service for data. The courier (Firehose) picks up packages (data records) and delivers them automatically to the address you specified (S3, Redshift, OpenSearch, etc.). You do not drive the truck yourself.
Because it is fully serverless, there is no infrastructure to manage — no shards, no servers. You simply send data and Firehose handles everything else.
Key features:
Automatic delivery destinations: Amazon S3, Amazon Redshift, Amazon OpenSearch Service, and HTTP endpoints are all supported out of the box. Buffering: Instead of delivering one record at a time, Firehose accumulates data until either a size threshold (e.g., 5 MB) or a time threshold (e.g., 60 seconds) is reached, then delivers a batch. This prevents thousands of tiny files from piling up in S3. Lambda transformation: Before delivering data, Firehose can invoke a Lambda function to transform it — for example converting JSON to Parquet, or masking sensitive fields. No custom consumers: Unlike Data Streams, Firehose only delivers to its predefined destinations. You cannot attach your own consumer applications. No replayability: Once data is delivered, you cannot re-read it from Firehose.
Kinesis Data Streams vs Kinesis Data Firehose:
| Aspect | Data Streams | Firehose | |--------|-------------|---------| | Management | Manual shard management | Fully serverless | | Consumers | Custom consumer apps | Fixed destinations only | | Replayability | Yes | No | | Latency | Milliseconds | Tens of seconds to minutes | | Best for | Complex real-time processing | Simple auto-loading to S3/Redshift |
Amazon MSK — Managed Kafka on AWS
Apache Kafka is an open-source streaming platform originally built by LinkedIn, widely used for large-scale real-time data streaming. Amazon MSK (Managed Streaming for Apache Kafka) is a service where AWS manages the Kafka cluster for you.
Think of the difference between "building and operating your own Kafka servers" versus "having AWS manage Kafka for you (MSK)." With MSK, AWS handles broker upgrades, patching, and failure recovery.
MSK vs Kinesis Data Streams:
| Aspect | Amazon MSK | Kinesis Data Streams | |--------|-----------|---------------------| | Underlying technology | Apache Kafka | AWS proprietary | | Ecosystem | Full Kafka ecosystem compatibility | AWS-native | | Migration | Existing Kafka code works as-is | Code rewrite needed | | Operational complexity | Relatively higher | Lower | | Choose when | You need the Kafka ecosystem | Starting a new AWS-native project |
Exam tip: If the question mentions "migrating an existing Kafka workload to AWS" or "Kafka ecosystem," choose MSK. If it says "AWS-native streaming," choose Kinesis Data Streams.
Batch Ingestion Services — Collecting and Processing in Bulk
For periodic or large-volume ingestion that does not require real-time processing, these services are the right tools.
AWS DMS (Database Migration Service)
DMS is like a professional moving company. It takes the furniture (data) from your old house (source database) and moves it to the new house (target database).
Homogeneous migration: MySQL to MySQL, Oracle to Oracle Heterogeneous migration: Oracle to Aurora PostgreSQL, SQL Server to MySQL CDC (Change Data Capture): DMS continuously captures every INSERT, UPDATE, and DELETE from the source database and applies those changes to the target. It is like following a database's "change diary" in real time. Full load plus CDC: First copy all existing data, then continuously sync only the changes going forward.
AWS AppFlow
AppFlow brings data from SaaS applications — Salesforce, SAP, Google Analytics, Slack, and more — into AWS. Instead of writing custom API integration code for each SaaS service, AppFlow lets you configure the connection with a few clicks.
Destinations: S3, Redshift, Salesforce, Snowflake, and others Runs on a schedule or triggered by events Can filter and transform data in transit No coding required for SaaS data ingestion
AWS Transfer Family
Transfer Family lets external partners upload and download files to S3 or EFS using SFTP, FTP, FTPS, or AS2 protocols. If a business partner sends you CSV files via SFTP every day and you want those files automatically stored in S3, Transfer Family is the answer.
S3 Event Notifications
When a file is uploaded to or deleted from an S3 bucket, S3 Event Notifications can automatically trigger another service. For example, when a partner uploads a CSV to your S3 bucket, a Lambda function immediately starts processing it. Notifications can be sent to SQS, SNS, Lambda, or EventBridge.
Exam Key Points Summary
| Keyword | Choose this service | |---------|-------------------| | Real-time streaming, custom processing, replayable | Kinesis Data Streams | | Real-time data auto-loaded to S3/Redshift, serverless | Kinesis Data Firehose | | Managed Apache Kafka, migrating existing Kafka workload | Amazon MSK | | Database migration, CDC | AWS DMS | | Ingest data from SaaS apps (Salesforce, etc.) | AWS AppFlow | | SFTP/FTP file transfer to S3 | AWS Transfer Family | | Auto-trigger processing when file uploaded to S3 | S3 Event Notifications + Lambda |
Data Streams vs Firehose decision guide: If you need custom consumer logic or the ability to replay data, choose Data Streams. If you simply need to automatically collect data into S3 or Redshift without custom processing, Firehose is the simpler and cheaper choice.