High-Performance Data Ingestion

Learn about Kinesis, Glue, Athena, Lake Formation for data ingestion, transformation, and analysis.

In the SAA-C03 exam, data pipeline questions make up roughly 16% of the total. Understanding whether to process data in real-time or in batches — and which service combinations are optimal — lets you confidently tackle this section.

 

What Is Real-Time Streaming?

Picture a factory conveyor belt. Products move along continuously and workers process each one without stopping. That is streaming processing. In contrast, collecting a day's worth of products in a warehouse and processing them all at night is batch processing.

In the exam, when you see "real-time", "immediately", or "continuous data flow", think streaming. When you see "nightly processing", "bulk data transformation", or "periodic execution", think batch.

 

Amazon Kinesis — The Core of Real-Time Streaming

Kinesis is a platform for collecting and processing large amounts of data in real time. It has three services.

 

Kinesis Data Streams

Collects data in real time and lets multiple consumers each process it in their own way. Like multiple teams on a factory floor each picking what they need from the same conveyor belt.

Collected data is retained for 24 hours by default (up to 365 days) Capacity is managed in shards: each shard handles 1 MB/s input and 2 MB/s output Multiple consumers (Lambda, Kinesis Data Analytics, EC2 apps) can read the same stream simultaneously You must implement your own custom processing logic

When to choose Kinesis Data Streams: look for "real-time ingestion + custom processing", "multiple consumers reading the same stream", "need to replay data".

 

Kinesis Data Firehose

Receives real-time data and automatically delivers it to destinations such as S3, Redshift, OpenSearch, or Splunk. Fully managed and serverless — no shard management needed.

Like an automatic sorter at the end of a conveyor belt that sends each product to the right warehouse on its own.

Buffers data before delivery (every 60 seconds or every 1–128 MB) Can attach a Lambda function to transform data before delivery No consumer code required — data lands at the destination automatically

When to choose Kinesis Data Firehose: look for "automatically save real-time data to S3/Redshift", "serverless streaming", "no custom consumer code needed".

 

Kinesis Data Analytics

Runs SQL or Apache Flink on streaming data for real-time analytics. Like watching items move along the conveyor belt and performing instant aggregation, filtering, or anomaly detection.

Useful for per-second aggregations, moving averages, and anomaly detection Output results back to Kinesis Data Streams or Firehose

 

Batch Processing and ETL

 

AWS Glue — Serverless ETL

ETL stands for Extract, Transform, Load. Glue automates this process without any server management.

Glue Crawlers: automatically scan S3, RDS, and other sources to detect schemas Glue Data Catalog: stores discovered schema metadata (referenced by Athena and EMR) Glue ETL Jobs: run Python or Scala scripts to transform data Glue Studio: visually build ETL pipelines without writing code

When to choose AWS Glue: look for "serverless ETL", "automatic schema discovery", "data catalog", "automate S3 data transformation".

 

Amazon EMR — Large-Scale Hadoop/Spark Clusters

EMR runs big data frameworks like Hadoop, Spark, Hive, and Presto on EC2 clusters. Use it for large-scale data transformation, machine learning preprocessing, and complex analytics.

Glue vs EMR: Glue is serverless (no management needed). EMR lets you configure clusters yourself for more flexible customization.

When to choose Amazon EMR: look for "Hadoop/Spark", "large-scale data transformation", "manage clusters directly", "custom big data frameworks".

 

AWS DataSync

Rapidly transfers large amounts of data from on-premises file systems to S3, EFS, or FSx. Use it for data migrations or regular synchronization tasks.

When to choose DataSync: look for "on-premises to S3 bulk data transfer" or "periodic file synchronization".

 

Analytics Services

 

Amazon Athena — SQL on S3

Analyzes data stored in S3 directly using SQL, without moving data to a separate database. Integrates with the Glue Data Catalog to automatically pull table definitions.

Fully serverless: you pay only for the data scanned by your queries Supports CSV, JSON, Parquet, ORC, and other formats Using columnar formats (Parquet) optimizes both query speed and cost

When to choose Athena: look for "query S3 data with SQL", "serverless queries", "analyze S3 data directly without a separate DB".

 

Amazon Redshift — Data Warehouse

A data warehouse for analyzing petabyte-scale structured data at high performance. Optimized for OLAP — large aggregation queries for analytics.

Redshift Spectrum: query S3 data directly from Redshift without moving it Automatic snapshots and auto-scaling available Remember: OLTP (transactions) = RDS/Aurora/DynamoDB. OLAP (analytics) = Redshift.

When to choose Redshift: look for "large-scale data warehouse", "OLAP", "large aggregation and analytics queries".

 

Amazon QuickSight — BI Visualization

A serverless BI tool that connects to Athena, Redshift, S3, and other sources to create dashboards and charts.

When to choose QuickSight: look for "BI dashboard" or "data visualization".

 

AWS Lake Formation — Data Lake Management

Makes it easy to build a data lake centered on S3 and manage fine-grained data access permissions. Integrates with Glue, Athena, and Redshift.

When to choose Lake Formation: look for "build a data lake", "central data access control", "column-level security".

 

Back to blog list